×
Register Here to Apply for Jobs or Post Jobs. X

Senior ML ​/ Evaluation Engineer

Job in Town of Poland, Jamestown, Chautauqua County, New York, 14701, USA
Listing for: Intellias
Full Time position
Listed on 2026-08-04
Job specializations:
  • Software Development
    AI QA / Validation Engineer, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 200000 USD Yearly USD 140000.00 200000.00 YEAR
Job Description & How to Apply Below
Location: Town of Poland

  • We are looking for a Senior ML / Evaluation Engineer to help define and implement quality standards for enterprise-grade AI agents and LLM-powered applications. In this role, you will design evaluation frameworks, build custom evaluation pipelines, and establish automated quality gates across the AI delivery lifecycle. You will work closely with AI Platform Engineers, ML Engineers, and Dev Ops teams to ensure reliable, measurable, and production-ready AI systems through scalable evaluation, observability, and governance practices.
  • Our customer is a multinational corporation with more than a century of history and offices in over 180 countries. Their most ambitious goal at the time is to introduce a range of Reduced-Risk Products (RRPs). The target audience is more than 1 billion consumers around the globe. IT platform hosts 700+ applications.
  • Intellia's mission is to help the client with the engineering of a comprehensive software ecosystem for a game-changing IoT product on the margin of innovative consumer experience and cutting-edge technology. Our teams are involved in the engineering of core platform components for best-in-class eCommerce, Digital Marketing and IoT solutions. As an Engineer, you will become a part of Core Architecture Team and be responsible for the architecture, implementation of best practices in our Digital Engineering Enterprise Platform.
  • The Platform is a set of services and internet applications that accelerate the development and delivery of software applications by taking care of common SDLC challenges. The Platform provides access and consumption for engineering teams to a set of services, technologies, practices for their development and for operating their application, ensuring a set of compliance and best practices.
Requirements:

Skills:
  • AWS Agent Core Evaluation (on-demand mode for CI/CD gates, online mode for production sampling)
  • Custom code-based Lambda evaluators (Python - deterministic checks)
  • Evaluation levels (TRACE for per-response,  for per-invocation, SESSION for workflow)
  • OTel spans from AWS Agent Core Observability as evaluation input
Experience:
  • 5+ years ML engineering or AI platform engineering
  • LLM evaluation framework design and implementation
  • Custom evaluator implementation for deterministic quality checks
Nice-to-have
  • AWS Agent Core Evaluation API hands-on (Create Evaluation, Get Evaluation Result )
  • Cloud Watch metrics output from Agent Core Evaluation for online mode
Responsibilities:
  • Design, implement, and maintain enterprise-grade evaluation frameworks for LLMs, AI agents, and multi-step AI workflows.
  • Develop and optimize LLM-as-a-judge evaluators to assess dimensions such as helpfulness, correctness, consistency, and policy compliance.
  • Build custom Python-based evaluators using AWS Lambda to perform deterministic validation, business-rule enforcement, and workflow quality checks.
  • Define and implement evaluation standards, mandatory quality dimensions, scoring methodologies, and pass/fail criteria across AI platforms.
  • Design evaluation strategies at multiple levels, including TRACE, , and SESSION evaluation scopes.
  • Integrate evaluation workflows into CI/CD pipelines and establish automated deployment quality gates for AI-powered applications.
  • Leverage AWS Agent Core Evaluation capabilities to execute on-demand evaluations and support production quality monitoring.
  • Utilize observability data, Open Telemetry traces, and Agent Core telemetry signals as evaluation inputs for quality assessment and root-cause analysis.
  • Collaborate with platform, security, and AI engineering teams to improve agent reliability, accuracy, and operational quality.
  • Analyze evaluation results, identify quality regressions, and drive corrective actions across models, prompts, tools, and workflows.
  • Define monitoring and reporting mechanisms for evaluation outcomes, quality trends, and operational KPIs.
  • Contribute to the evolution of enterprise AI governance, testing methodologies, and evaluation best practices.
#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary