×
Register Here to Apply for Jobs or Post Jobs. X

Machine Learning Engineer - Agentic AI Evaluation Frameworks

Job in Northern, Floyd County, Kentucky, USA
Listing for: Apple Inc.
Full Time position
Listed on 2026-09-10
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 184700 - 277600 USD Yearly USD 184700.00 277600.00 YEAR
Job Description & How to Apply Below
Location: Northern

Machine Learning Engineer - Agentic AI Evaluation Frameworks

Cupertino, California, United States Machine Learning and AI

Imagine what you could do here. At Apple, great ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your work, and there’s no telling what you could accomplish.

The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered  this role, you will develop evaluation frameworks, datasets, tooling, and quality signals that enable teams to understand and continuously improve Generative AI and LLM-powered products. You will work closely with Machine Learning, Software Engineering, Quality Engineering, Product, Human Interface, Data Science, and domain experts to establish rigorous evaluation practices throughout the AI product lifecycle.

You will help define how we measure the quality of AI experiences across the Commerce domain, including Store AI, Shopping AI, Learning AI, Content GenAI, Conversational AI, and Platform Self-Service.

This is an opportunity to work at the intersection of machine learning, software engineering, data, and product quality, helping ensure our AI experiences are accurate, relevant, grounded, reliable, and useful for users around the world.

Description

As a Machine Learning Evaluation Engineer, you will design and build scalable evaluation systems for LLM, Generative AI, Conversational AI, and Agentic AI products.

  • Design and develop automated evaluation frameworks and pipelines for AI-powered products.
  • Define evaluation methodologies and quality metrics across dimensions such as accuracy, relevance, groundedness, completeness, consistency, instruction following, and task completion.
  • Build and maintain high-quality evaluation datasets, including golden datasets, benchmark sets, regression suites, adversarial scenarios, and production-derived test sets.
  • Develop Auto Eval capabilities that enable teams to rapidly evaluate models, prompts, retrieval systems, agents, and end-to-end AI experiences.
  • Design and implement model-based evaluation approaches, including LLM-as-a-Judge, while developing appropriate calibration and validation methodologies.
  • Develop Human-in-the-Loop (HITL) evaluation approaches for complex or subjective quality dimensions where automated evaluation alone is insufficient.
  • Define evaluation rubrics, annotation guidelines, grading criteria, and quality standards in partnership with product teams, domain experts, and annotation teams.
  • Build mechanisms to calibrate automated evaluators against human judgment and measure evaluator consistency and reliability.
  • Evaluate end-to-end AI systems, including retrieval, context construction, prompts, model responses, tool use, APIs, and downstream product experiences.
  • Develop evaluation methodologies for multi-turn conversations, personalization, recommendations, tool use, reasoning, and agentic task execution.
  • Perform detailed error analysis and failure-mode investigation to identify opportunities for model, prompt, retrieval, dataset, and product improvements.
  • Build reusable evaluation infrastructure, APIs, dashboards, and developer tooling that can scale across multiple AI products and teams.
  • Integrate evaluation into development and CI/CD workflows, enabling automated regression detection, quality gates, and release-readiness assessments.
  • Connect offline evaluation results with production signals to continuously improve evaluation coverage and product quality.
  • Partner closely with Machine Learning, Software Engineering, Product, Quality Engineering, Human Interface, and Data Science teams throughout research, development, evaluation, launch, and continuous improvement.
Minimum Qualifications
  • Typically requires a minimum of 7 years of related experience in Machine Learning Engineering, ML Evaluation, Software Engineering, Data Science, Quality Engineering, or a related technical field.
  • Strong programming skills in Python and experience developing production-quality software, ML systems, data pipelines, or evaluation infrastructure.
  • Experience developing or evaluating LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or other machine-learning-driven products.
  • Experience designing automated ML evaluation frameworks, metrics, benchmarks, datasets, or experimentation methodologies.
  • Understanding of modern LLM application…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary