Agent Evaluation Engineer — Build-Time Framework & Deployment Gates
Remote / Online - Candidates ideally in
Jacksonville, Duval County, Florida, 32290, USA
Listed on 2026-10-02
Jacksonville, Duval County, Florida, 32290, USA
Listing for:
EPAM Systems Inc
Remote/Work from Home
position Listed on 2026-10-02
Job specializations:
-
Software Development
AI QA / Validation Engineer, AI Reliability/ Performance Engineer
Job Description & How to Apply Below
We're looking for an Agent Evaluation Engineer — Build-Time Framework & Deployment Gates to join our team in Portugal in a fully remote working mode. In this role, you will design and maintain an evaluation framework for AI agents, ensuring quality and compliance through automated tests and CI/CD deployment gates. You will develop multi-layer evaluation suites that blend deterministic checks with LLM-powered graders, simulate multi-turn conversations, and define reliability metrics.
The position also involves implementing staging validations, shadow-mode traffic analysis, and A/B rollout strategies, with feedback loops from production environments to enhance overall system robustness.
- Design and implement build-time evaluation frameworks for agentic workflows using Lang Graph or comparable orchestration frameworks
- Create deterministic and LLM-as-judge grading pipelines covering reasoning, trajectory accuracy, and output quality
- Develop test harnesses for multi-turn conversational simulations and context-retention scoring
- Define reliability assessment methods including multi-trial metrics (pass@k, pass^k)
- Implement CI/CD deployment gates that enforce quality thresholds and block releases not meeting standards
- Integrate staging validation, shadow-mode traffic comparison, and A/B rollout control in deployment pipelines
- Leverage AWS Agent Core Evaluations for on-demand and online scoring components connected to production feedback
- Convert production incidents into reusable regression cases for continuous quality improvement
- Collaborate with engineering and Dev Ops teams to embed evaluation gates into automated workflows
- 4+ years of experience in building automated testing or evaluation frameworks for ML, LLM, or agentic systems
- Proven hands‑on experience designing multi-layer evaluation suites with deterministic and LLM-based graders
- Expertise with CI/CD pipelines and implementing metric-based quality gates for automated deployments
- Practical knowledge of Lang Graph or similar agent orchestration frameworks
- Strong background in designing simulation-based evaluation strategies and conversation-level tests
- Nice to have Experience with AWS Agent Core Evaluations API (Create Evaluation, custom evaluators)
- Familiarity with shadow-mode, canary, or A/B deployment practices for ML-based platforms
- Background in transforming production failures into build-time regression tests for agent workflows
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×