Principal Evals Engineer - AI & Agentic Systems
Listed on 2026-09-22
-
Software Development
AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
Principal Evals Engineer - AI & Agentic Systems
About the Organisation
Our client is building one of the region’s most ambitious government AI programmes, developing the AI infrastructure, applications and platforms required to operate AI-native public services at scale.
The systems being built must operate reliably across Arabic and English
, within strict requirements around data sovereignty, privacy, security and regulated infrastructure.
These are production systems supporting consequential workflows and senior decision-makers. Whether an AI system works cannot simply be a matter of opinion.
It has to be measurable.
This role owns how that measurement happens.
Why Join- Solve difficult problems in applied AI. Determine whether non-deterministic AI systems actually work across real-world, bilingual and high-stakes environments.
- Own the quality bar. Define the evidence required before model changes, prompt updates, framework migrations or architectural changes reach production.
- Small, senior engineering teams. Work alongside experienced AI, product and forward-deployed engineers with significant technical autonomy.
- Build the evaluation practice. Establish the evaluation architecture, tooling and engineering standards that teams across the programme will use.
- High-impact environment. Your work will influence AI systems deployed across major public-sector organisations and used at significant scale.
- Relocation supported. Relocation assistance is available for successful candidates and their families.
We are hiring a Principal Evals Engineer to own how the organisation determines whether its AI systems actually work.
You will set the direction for evaluation across a portfolio spanning AI assistants, retrieval systems, agentic workflows, voice applications and document intelligence
, while building the infrastructure that turns “it seems better” into measurable evidence.
You own measurement and release evidence
, partnering closely with Principal-level engineers responsible for AI architecture, product engineering, platform development and forward deployment.
Evaluating AI systems is fundamentally different from conventional software testing. The same input can produce different outputs, correctness can be subjective, and failures such as hallucination, poor grounding, behavioural drift and prompt injection cannot be captured through conventional assertions alone.
At Principal level, you will be expected to have already tackled these problems in production environments.
This is primarily an individual contributor role
. Your influence comes from the infrastructure you build, the quality bar you establish and the engineering decisions your evidence enables.
Evaluation strategy and architecture.
Define what gets measured, at which layer, using which methodology and how results feed into product and release decisions. Establish the reference architecture engineering teams build against.
The shared evaluation platform.
Build evaluation harnesses, golden-set management, dataset versioning, automated grading, behavioural regression detection and reporting infrastructure.
Grading you can trust.
Own judge-model selection, rubric design and calibration against human labels — including understanding when automated grading cannot be trusted.
Retrieval, agents and multilingual evaluation.
Measure grounding, citation correctness, tool use, multi-step reasoning and failure recovery. Build dedicated Arabic evaluation datasets and ensure judges are properly calibrated rather than assuming English evaluation methods transfer directly.
Evidence behind engineering decisions.
Model swaps, prompt changes, framework migrations and infrastructure decisions should be supported by defensible evaluation evidence before…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).