×
Register Here to Apply for Jobs or Post Jobs. X

Machine Learning Platform Engineer, AI Evaluation

Job in Seattle, King County, Washington, 98127, USA
Listing for: Apple Inc.
Full Time position
Listed on 2026-07-18
Job specializations:
  • Software Development
    AI Reliability/ Performance Engineer, Backend Developer, AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 205400 - 308500 USD Yearly USD 205400.00 308500.00 YEAR
Job Description & How to Apply Below
Position: Staff Machine Learning Platform Engineer, AI Evaluation

Staff Machine Learning Platform Engineer, AI Evaluation

Seattle, Washington, United States Machine Learning and AI

Join Apple Services Engineering to build the next generation of AI evaluation systems. We are seeking a staff machine learning platform engineer to lead the architectural design and development of the high availability services and internal tools powering self‑service evaluation  will partner with researchers to operationalize their innovations, transforming complex workflows into intuitive, developer‑first platforms. We are looking for builders who thrive in the ambiguity of new initiatives and are passionate about creating scalable infrastructure.

Description

We're building the evaluation platform that will serve all of Apple's generative AI and agent systems. This is early‑stage work – some scrappy components exist, much is greenfield and we need a staff engineer who can take it from here to org‑wide self‑service scale. This is not a “maintain the infra” role. You'll make consequential decisions about what to build, what to integrate, and what to say no to, then ship it in Python with a small team.

Responsibilities
  • Platform architecture & delivery:
    Own the technical direction for our evaluation platform. Design and build the APIs, SDKs, and orchestration services that turn research‑grade evaluation methodology into self‑service building blocks other teams ship on top of. You'll start scrappy and intentionally, with line of sight to the scaled version.
  • Productionize ML research:
    Partner directly with research engineers to assess their code and determine what can be rewritten into clean Python services vs. what requires infrastructure changes (Ray, GPU compute, distributed scheduling). Build the reusable abstractions that make the next research handoff faster than the last.
  • Strategic decision‑making:
    You will balance complex, competing priorities from partner engineering teams, PMs, and leadership. Your job is to distinguish signal from noise, identifying the platform‑level decisions that serve the org vs. one‑off requests that don't scale. You'll advocate for these decisions clearly in documentation and in rooms with senior stakeholders.
  • Drive org‑level evaluation strategy:
    Work with your technical manager to assess workload across engineers, set priorities, and define how self‑service evaluation reaches every team 're a force multiplier, not just through code, but through the clarity of your technical vision.
  • Developer experience:
    You own the experience end‑to‑end. Today that means supporting existing evaluation patterns (trace‑based, metric‑based). Tomorrow it means enabling breakthrough approaches — surfacing where models fail in non‑obvious ways, evaluating multi‑turn agent trajectories, and scoring complex tool‑use chains.
  • Operational rigor:
    Define the team's posture on testing, CI/CD, monitoring, and reliability. You don't need to be an SRE, but you ship with instrumentation and you set the standard others follow.
Minimum Qualifications
  • 8+ years of software engineering experience with a track record of owning platform‑level technical direction.
  • 0‑to‑1 builder who designs for scale. You've taken something from nothing to production, made deliberate tradeoffs about what to build now vs. later, and can articulate why.
  • ML depth:
    You're not building the models, but you can read research code and assess: is this a software problem or an infrastructure problem? Do we need a rewrite or do we need GPUs? You speak the language of research engineers fluently.
  • AI/Agent evaluation experience that goes beyond traces. You understand the hard problems: non‑deterministic outputs, multi‑step agent reasoning, judge model reliability, scoring drift. You've built or operated systems that handle these.
  • Judgment under ambiguity. You know when to build a rapid prototype for quick validation and when to be disciplined (design doc, review, test). You can tell the difference in real time, not just in retrospect.
  • Communication as a core skill. You write clearly design docs, decision records, platform roadmaps. You speak clearly in meetings with researchers, in rooms with engineering leaders, and balance the…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary