Principal Evals Engineer
Listed on 2026-10-02
-
Software Development
AI QA / Validation Engineer, AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
We are hiring a Principal Evals Engineer to join a major new AI programme in Abu Dhabi. The organization is building a large scale AI platform designed to support next generation AI systems across multiple sectors. The work spans AI assistants, retrieval systems, agentic workflows, voice applications and document intelligence, with a strong focus on building reliable AI systems that can operate at scale in real world environments.
Discoverthe Role
- As Principal Evals Engineer, you will define how AI performance is measured across the organization.
- You will design and build the shared evaluation architecture used across multiple AI products and engineering teams. This includes creating robust evaluation harnesses, managing golden datasets, automating grading, detecting behavioural regressions and ensuring that engineering teams can make release decisions based on evidence rather than subjective judgement.
- This is a senior individual contributor role. You will remain highly hands on, building frameworks, running experiments and setting the technical standard for AI evaluation across the organization.
- Define the overall AI evaluation strategy and reference architecture across the engineering organization.
- Build shared evaluation frameworks, harnesses and reusable libraries.
- Develop and maintain golden datasets and dataset versioning strategies.
- Design automated evaluation and grading systems.
- Implement regression detection for non deterministic AI behaviour.
- Evaluate LLM output quality, hallucination, grounding, faithfulness and citation accuracy.
- Build evaluation frameworks for RAG and retrieval systems.
- Evaluate agentic systems, including tool usage, multi step reasoning and failure recovery.
- Design and calibrate LLM as a judge systems against human evaluation.
- Establish production monitoring, online evaluation and drift detection.
- Develop adversarial testing methodologies covering prompt injection, jailbreaks and data leakage.
- Integrate AI evaluation into CI/CD pipelines and release processes.
- Help engineering teams understand and apply rigorous evaluation standards independently.
- Contribute to building the broader evaluation engineering function as the organization scales.
We are looking for a senior engineer with a strong track record of building AI or testing infrastructure at scale.
You should have experience owning evaluation frameworks rather than simply working inside an existing testing platform.
Discover the Requirements- Staff or Principal level engineering experience.
- Strong hands on Python engineering skills.
- Deep experience with LLM or Generative AI evaluation.
- Experience evaluating RAG, retrieval or agentic systems.
- Strong understanding of hallucination, grounding and output quality measurement.
- Experience building regression detection for non deterministic AI systems.
- Practical experience with LLM as a judge techniques.
- Strong understanding of rubric design and calibration against human labels.
- Statistical literacy including sampling, confidence intervals, significance and inter rater agreement.
- Strong CI/CD and software quality engineering experience.
- Experience building reusable frameworks or shared engineering libraries.
- Hands on experience working with production AI systems.
Experience with one or more of the following would be highly valuable:
- Ragas.
- Deep Eval.
- Promptfoo.
- Braintrust.
- Lang Smith.
- Langfuse.
- Lang Chain or Lang Graph.
- Vector databases such as pgvector or Qdrant.
- AI red teaming or adversarial testing.
- Voice or conversational AI evaluation.
- Human evaluation programme design.
- Data pipeline or performance testing.
- Arabic language AI evaluation, including golden datasets, dialect coverage and right to left validation.
- Experience working within regulated, government or highly secure environments.
The primary language is Python, with additional use of Type Script and Java.
The wider environment includes custom evaluation harnesses, LLM as a judge patterns, golden datasets, Langfuse, Ragas, Deep Eval, Promptfoo, Lang Chain, Lang Graph, Microsoft Agent Framework, pgvector, Qdrant, Playwright, Cypress, Great Expectations, Grafana, Docker, Kubernetes and Azure.
The team is pragmatic about tooling and is more interested in strong engineering judgement than experience with one specific framework.
Why join?This is an opportunity to help shape the evaluation discipline within one of the most ambitious AI programmes currently being built in the…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).