×
Register Here to Apply for Jobs or Post Jobs. X

Senior Research Engineer, Training Data Infrastructure in Foundation Models

Job in Cupertino, Santa Clara County, California, 95014, USA
Listing for: Apple Inc.
Apprenticeship/Internship position
Listed on 2026-07-18
Job specializations:
  • Software Development
    Data Engineering
Salary/Wage Range or Industry Benchmark: 184700 - 324800 USD Yearly USD 184700.00 324800.00 YEAR
Job Description & How to Apply Below

Senior Research Engineer, Training Data Infrastructure in Foundation Models

Cupertino, California, United States Software and Services

We build frontier foundation models that power intelligent experiences  team works across the full training lifecycle: including pre‑training foundation models, and developing mid‑training approaches that bridge general capability and task‑specific performance. What makes our work distinct is that we’re engineering models specifically for Apple silicon and optimized for experiences that are private, personal, and deeply integrated into the OS. We’re solving frontier problems in reward modeling to resist reward hacking, handling sparse and delayed rewards in agentic settings, and aligning models reliably across the spectrum from open‑ended creative tasks to precise, action‑taking workflows.

If you’re drawn to hard problems where the research and the product are inseparable, this is the team.

Description

This position operates at the convergence of Software Engineering and Machine Learning Research. Unlike traditional backend roles, this position requires you to design systems where the outcome is the statistical distribution and quality of data itself. You will work alongside Research Scientists to transform theoretical observations into concrete, scalable engineering solutions. Your core focus will be the architecture of our Data Acquisition, Processing, and Repository Management systems for Large Model training.

You will lead technical efforts to enable active, quality‑driven data curation, including filtering, deduping, synthetic data generation and data mixing, ensuring our models are trained on the highest‑quality information available.

Responsibilities
  • Architect Scalable Ingestion Systems:
    Design and implement high‑throughput distributed systems to ingest petabytes of text and multimodal data from diverse sources, including web crawls and third‑party partnerships.
  • Repository Optimization:
    Manage the lifecycle of large‑scale datasets across data storage and high‑performance file systems. Optimize data formats for efficient random access and sequential scanning during model training.
  • Data Governance & Privacy:
    Engineer robust data governance and privacy solutions for the training data, in collaboration with compliance and legal teams, to ensure adherence to stringent regulatory standards.
  • High‑Performance Processing Pipelines:
    Build and maintain distributed data processing workflows using advanced frameworks on cloud infrastructure (e.g., GCP, AWS).
  • Algorithmic Data Curation:
    Implement sophisticated data filtering and selection logic to remove low‑quality content. Develop semantic deduplication at scale to prevent model memorization and improve training efficiency.
  • Decontamination Removal:
    Design automated systems to detect and remove benchmark leakage, ensuring that evaluation datasets remain strictly isolated from training corpora.
  • Infrastructure for Scaling Laws:
    Collaborate with researchers to enable data ablations and scaling experiments. Build tools to support systematic data mixture optimization and empirically data studies.
Minimum Qualifications
  • Education:

    Bachelor’s degree in Computer Science, Electrical Engineering, or Mathematics.
  • Technical Expertise: 4+ years of software engineering experience with a specific focus on Data Infrastructure, Distributed Systems, or AI/ML Engineering.
  • Language Proficiency:
    Expert fluency in Python, and strong competence in system languages such as C++.
  • Cloud Architecture:
    Extensive experience architecting solutions on major public cloud platforms (e.g., GCP) to build scalable data systems (e.g., with Apache Beam, GCS).
  • Performance Engineering:
    Deep experience profiling and optimizing high‑throughput data systems. Demonstrated ability to debug distributed bottlenecks (e.g., stragglers, I/O saturation), optimize data formats and provide efficient data storage solutions.
Preferred Qualifications
  • Research

    Collaboration:

    Experience working within or closely with ML research organizations (e.g., as a Research Engineer), with an ability to translate research results into engineering implementations.
  • Domain Knowledge:
    Familiarity with…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary