×
Register Here to Apply for Jobs or Post Jobs. X

Agent Evaluation Engineer — Build-Time Framework & Deployment Gates

Remote / Online - Candidates ideally in
Jacksonville, Duval County, Florida, 32290, USA
Listing for: EPAM Systems Inc
Remote/Work from Home position
Listed on 2026-10-02
Job specializations:
  • Software Development
    AI QA / Validation Engineer, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below

We're looking for an Agent Evaluation Engineer — Build-Time Framework & Deployment Gates to join our team in Portugal in a fully remote working mode. In this role, you will design and maintain an evaluation framework for AI agents, ensuring quality and compliance through automated tests and CI/CD deployment gates. You will develop multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulate multi-turn conversations, and define reliability metrics.

The position also involves implementing staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production environments to enhance overall system robustness.

Responsibilities
  • Design and implement build-time evaluation frameworks for agentic workflows using Lang Graph or comparable orchestration frameworks
  • Create deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output quality
  • Develop test harnesses for multi-turn conversational simulations and context-retention scoring
  • Define reliability assessment methods including multi-trial metrics (pass@k, pass^k)
  • Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standards
  • Integrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelines
  • Leverage AWS Agent Core Evaluations for on-demand and online scoring components connected to production feedback
  • Convert production incidents into reusable regression cases for continuous quality improvement
  • Collaborate with engineering and Dev Ops teams to embed evaluation gates into automated workflows
Requirements
  • 4+ years of experience in building automated testing or evaluation frameworks for ML, LLM, or agentic systems
  • Proven hands‑on experience designing multi-layer evaluation suites with deterministic and LLM-based graders
  • Expertise with CI/CD pipelines and implementing metric-based quality gates for automated deployments
  • Practical knowledge of Lang Graph or similar agent orchestration frameworks
  • Strong background in designing simulation-based evaluation strategies and conversation-level tests
  • Nice to have Experience with AWS Agent Core Evaluations API (Create Evaluation, custom evaluators)
  • Familiarity with shadow-mode, canary, or A/B deployment practices for ML-based platforms
  • Background in transforming production failures into build-time regression tests for agent workflows
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary