Senior AI Agent Quality Engineer
Listed on 2026-09-08
-
Software Development
AI QA / Validation Engineer, AI Engineer (Applied/Software)
We are seeking a Senior AI Operations Engineer to serve as the quality gatekeeper for the agentic AI systems powering Claritev's next generation of healthcare products.
This is a hands-on Agentic QA role for an engineer who is passionate about answering a deceptively hard question: how do you prove an AI agent works? You will design and run the test suites, benchmarks, and evaluation pipelines that validate agent behavior before and after release - catching regressions, hallucinations, broken tool calls, and unsafe actions before they reach production healthcare workflows.
You will work closely with our Principal Agentic AI Operations Engineer, AI engineering teams, and Product stakeholders to build golden datasets, automate agent testing in CI/CD, triage and reproduce agent failures, and turn evaluation results into actionable quality insights. This role is an excellent path for a strong QA/test automation, MLOps, or backend engineer looking to specialize in the fast-growing discipline of AI agent quality and evaluation.
Job Role s and Responsibilities- Design, build, and maintain automated test suites and evaluation pipelines for agentic AI systems, covering single-turn, multi-turn, and end-to-end workflow scenarios.
- Develop and curate golden datasets, test scenarios, and simulation environments that reflect real healthcare workflows and edge cases.
- Execute and automate agent benchmarking, including regression testing across model, prompt, and tool changes, with clear pass/fail quality gates in CI/CD.
- Measure and report agent quality metrics: task completion, tool-call accuracy, grounding/hallucination rates, response quality, latency, cost per task, and safety compliance.
- Implement LLM-as-judge and rubric-based scoring approaches, and validate automated scores against human review.
- Triage, reproduce, and root-cause agent failures by analyzing traces of LLM calls, tool invocations, retrieval steps, and orchestration paths; file actionable defect reports.
- Perform adversarial and negative testing, including prompt-injection probes, malformed inputs, boundary conditions, and guardrail verification.
- Monitor production agent behavior, investigate quality incidents and drift, and feed findings back into offline test coverage.
- Contribute to evaluation dashboards, runbooks, and quality documentation used by engineering and product teams.
- Verify PHI/PII handling, auditability, and human-in-the-loop controls behave as designed, supporting HIPAA and data-governance compliance.
- Collaborate with AI engineers and data scientists to make agents more testable, observable, and reliable by design.
- Bachelor's degree in Computer Science, Engineering, Data Science, a quantitative discipline, or a related field required.
- Master's degree a plus.
- 5+ years of hands-on experience in software engineering, QA/test automation engineering, ML engineering, or a related technical discipline.
- 2+ years of experience testing, operating, or supporting production ML or AI systems.
- 1+ years of hands-on experience with generative AI, LLMs, RAG, and/or agentic AI systems (professional projects preferred; substantial personal or open-source work considered).
- Track record of building test automation or evaluation capabilities that measurably improved software or model quality.
- Strong Python skills, including test frameworks (e.g., pytest) and building automation/tooling; comfort working with APIs and distributed systems.
- Hands-on experience with LLM/agent evaluation concepts: eval datasets, LLM-as-judge, rubric-based scoring, and regression detection.
- Familiarity with evaluation and observability tooling such as Lang Smith, Langfuse, Arize Phoenix, Ragas, Deep Eval, promptfoo, or equivalent.
- Working knowledge of agentic AI frameworks and patterns (e.g., Lang Graph, Lang Chain, Auto Gen, CrewAI), including tool use, orchestration, and guardrails.
- Understanding of RAG pipelines, embeddings, and vector search, and how retrieval quality affects agent behavior.
- Experience with CI/CD pipelines and integrating automated tests and quality gates into them.
- Familiarity with cloud environments, containers, and log/trace analysis for debugging distributed systems.
- Solid grasp of statistics fundamentals for interpreting evaluation results (sampling, variance, significance).
- Strong analytical and debugging skills, with high attention to detail and a healthy skepticism toward "it works on my machine."
- Clear written communication, especially in defect reports, test plans, and quality summaries.
- Ability to operate effectively in a fast-moving, cross-functional environment.
- Experience with Oracle Cloud Infrastructure (OCI), including its generative AI, data science, database, and observability capabilities.
- Experience in healthcare, health technology, insurance, claims, payment integrity, or other regulated industries.
- Experience testing systems that process sensitive data, including PHI or PII, in…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).