Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Listed on 2026-08-30
-
IT/Tech
Machine Learning/ ML Engineer, AI Evaluation, AI Engineer (Applied/Software)
Machine Learning Engineer, Human Centered AI - Evaluations & Insights
Seattle, Washington, United States Machine Learning and AI
Imagine what you could do here. At Apple, great new ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your job and there's no telling what you could accomplish! Are you passionate about music, movies, and the world of Artificial Intelligence and Machine Learning? So are we!
Join our Human-Centered AI team for Apple Media Services. In this role, you'll represent the user perspective on new features, review and analyze data, and evaluate AI models powering everything from search and recommendations to other innovative features. You'll also collaborate with Data Scientists, Researchers, and Engineers to drive improvements across our platforms.
We are looking for a Machine Learning Engineer focused on Evaluation & Insights for the Human-Centered AI team. In this role, you will bridge the gap between human perception and algorithmic performance, helping evaluate and optimize Foundation Models and generative AI systems. You will architect robust evaluation frameworks, design scalable MLOps pipelines for model assessment, and translate qualitative failure modes into programmatic guardrails and training signals (e.g., SFT, RLHF/DPO).This
role blends deep ML engineering expertise with strong analytical judgment to assess, interpret, and improve the behavior of advanced AI models. You will work cross-functionally with Software Engineering, Product, Research and Responsible AI teams at Apple to ensure that our AI experiences are reliable, safe, and aligned with human expectations.
- Lead Rigorous Model Evaluations:
Architect and execute comprehensive evaluation suites for LLMs and multimodal models, identifying edge cases in multi-step reasoning, factuality, adversarial robustness, safety, and alignment. - Advanced Scoring Frameworks:
Develop deterministic, heuristic, and LLM-assisted evaluation frameworks (e.g., LLM-as-a-judge, reward modeling) to quantify human-perceived quality metrics (e.g., helpfulness, hallucination rates). - Actionable Signal Extraction:
Translate qualitative failure modes into quantifiable loss patterns, programmatic guardrails, and actionable data-mixture adjustments for model training and inference. - Improve Performance:
Partner with engineering teams to refine model behavior, leveraging evaluation telemetry to inform prompt engineering, Retrieval-Augmented Generation (RAG) strategies, and model fine-tuning. - Latent Pattern Recognition:
Apply advanced ML techniques (e.g., embedding-based clustering, representation learning, perturbation analysis) to systematically map error taxonomies and latent failure manifolds in model outputs. - MLOps & Automation:
Develop robust MLOps workflows to codify evaluation metrics, automate regression testing across model checkpoints, and integrate human-centric assessments into ML CI/CD pipelines. - Distributed Evaluation Pipelines:
Architect scalable, distributed inference and processing pipelines (e.g., Ray, vLLM) for high-throughput model evaluation, automated annotation, and output analysis at scale. - Human-Centric Metrics:
Define quantitative evaluation frameworks that capture nuanced human factors, including trust calibration, conversational state tracking, and interpretability. - Auto-Evaluator Systems:
Build automated evaluation pipelines utilizing LLMs to assess outputs at scale, optimizing for high correlation with human baseline annotations. - Cross-Functional Partnership:
Collaborate with ML researchers, software developers, and product managers across Apple to translate product requirements into scalable, reliable, and efficient model evaluation infrastructure.
- 5+ years of relevant industry experience in ML Engineering or Applied Research.
- Advanced proficiency in Python and modern deep learning ecosystems (PyTorch, JAX, Hugging Face).
- Proven experience building scalable ML inference pipelines, model-evaluation workflows, and structured rating frameworks for large-scale AI systems.
- Strong ability to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).