Senior Machine Learning Engineer, AI Evaluation
Listed on 2026-08-27
-
IT/Tech
AI Evaluation, Machine Learning/ ML Engineer
SHRMis a member-driven catalyst for creating better workplaces where people and businesses thrive together.
As the trusted authority on all things work, SHRM is the foremost expert, researcher, advocate, and thoughtleader on issuesand innovations impacting today’s evolving workplaces.
With nearly
340,000members in180countries, SHRM touches the lives of more than
362million workers and their families globally.
The Senior Machine Learning Engineer, AI Evaluation builds and operates the measurement and engineering infrastructure supporting the organization's Applied AI Research (AAIR) function, a continuous experimental environment designed to evaluate how artificial intelligence models perform real-world HR and workplace-related tasks against established professional standards.
This role is responsible for designing and maintaining the engineering infrastructure used to conduct rigorous, reproducible AI model evaluations and benchmarks. The Senior Machine Learning Engineer develops the systems that run multiple AI models against structured, domain-specific evaluations; builds scoring and evaluation frameworks; maintains reproducibility across model versions; and creates the data infrastructure necessary to analyze and track model performance over time.
Working closely with HR subject matter experts and Applied AI Research colleagues, this position translates expert-defined standards and evaluation criteria into technically rigorous, measurable specifications. Subject matter experts establish the domain-specific ground truth and standards for correctness, while the Senior Machine Learning Engineer owns the technical systems and methodologies used to measure model performance against those standards.
The position serves as a shared technical engineering resource across multiple Applied AI Research work streams and helps ensure that published findings, benchmarks, and research conclusions are supported by reliable, auditable, and defensible measurement practices.
This is an AI evaluation and engineering infrastructure role rather than a model-training or frontier AI research position. The role is not responsible for developing novel model architectures or training foundation models.
Responsibilities AI Evaluation Engineering & Infrastructure- Design, build, and maintain scalable engineering infrastructure for conducting structured evaluations and experiments across multiple AI and large language model (LLM) families.
- Develop and maintain a unified, provider-agnostic orchestration layer that enables consistent evaluation across multiple frontier model providers and architectures.
- Design and implement rigorous AI evaluation and scoring frameworks, including rubric-based scoring, model-as-judge methodologies with appropriate safeguards, partial-credit methodologies, and approaches for managing ambiguity.
- Build systems and processes that support reproducible experimentation, including model-version pinning, comprehensive run logging, experiment tracking, and drift detection.
- Maintain portable evaluation architecture across AI model providers to enable consistent and defensible cross-model comparisons as models and technologies evolve.
- Establish and maintain technical standards and engineering practices that support reliable, repeatable, and auditable AI evaluation.
- Partner closely with HR subject matter experts to translate professional standards, research criteria, and judgment-based rubrics into measurable and technically executable evaluation specifications.
- Identify and surface ambiguity, inconsistencies, or measurement limitations within proposed evaluation criteria and collaborate with subject matter experts to strengthen evaluation design.
- Apply knowledge of AI evaluation methodologies, benchmarking techniques, inter-rater reliability, and known limitations of automated and model-as-judge evaluation approaches.
- Support the design of measurement methodologies when definitive ground truth is unavailable or requires expert interpretation.
- Ensure evaluation methodologies align with research-defined validation standards and produce findings that are reproducible,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).