Senior AI Evaluation Engineer LangGraph
Listed on 2026-09-04
-
Software Development
AI QA / Validation Engineer
Job Title
Job Description
ResponsibilitiesDesign and develop build time evaluation frameworks for Lang Graph based agent systems
Create automated test harnesses for graph level and node level validation
Design evaluation strategies that combine deterministic grading and LLM as judge methodologies
Implement multi layer evaluation frameworks covering tool selection accuracy, execution trajectory quality, reasoning effectiveness, and output quality
Develop multi turn conversation simulations and context retention scoring mechanisms
Implement multi trial reliability testing methodologies, including pass at k and pass power k approaches
Design and maintain CI/CD deployment gates that validate quality metrics before production releases
Build staging validation, shadow mode comparison, and controlled rollout evaluation workflows
Integrate production evaluation feedback into build time testing frameworks to improve quality coverage
Collaborate with platform and engineering teams to establish evaluation standards, thresholds, and governance practices
Analyze evaluation results and provide recommendations for improving agent reliability and performance
Contribute to technical documentation, testing standards, and quality engineering best practices
Requirements4+ years of experience building automated testing frameworks, evaluation platforms, or quality assurance solutions for machine learning, large language model, or agent based systems
Hands on experience designing multi layer evaluation frameworks that combine deterministic and LLM as judge grading approaches
Experience implementing automated quality gates that can block deployments based on predefined metric thresholds
Experience working with Lang Graph or a comparable agent orchestration framework
Strong understanding of agent behavior evaluation, workflow validation, and AI quality measurement techniques
Experience designing scalable testing and validation processes for production AI systems
Strong Python development experience
Knowledge of CI/CD practices, deployment automation, and release governance
Experience analyzing evaluation data and translating findings into platform improvements
Strong communication and collaboration skills
Nice to HaveHands on experience with AWS Agent Core Evaluations, including evaluation execution and custom evaluator development
Experience designing and operating shadow mode or canary deployment strategies for machine learning or AI systems
Experience creating automated feedback loops that convert production incidents into regression test scenarios
Knowledge of production observability, monitoring, and evaluation pipelines
Experience with enterprise AI governance and quality assurance programs
Understanding of agent observability and telemetry driven quality improvement processes
We OfferVacation as per the laws of your country
Health insurance to help you take out an insurance policy for you and your loved ones
Sick pay: 10 days without a doctor's note, afterwards - as per the laws of your country
Equal opportunities
Time off for state holidays according to the official calendar, regardless of the client's schedule
Pleasant environment with two large corporate parties and many small get-togethers for colleagues
Comfort service solving technical and everyday problems at work
More about the perks:
List of benefits in PDF. The benefits package may vary depending on the region and the type of contract
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).