×
Register Here to Apply for Jobs or Post Jobs. X

Principal Evals Engineer - AI & Agentic Systems

Job in Abu Dhabi, UAE/Dubai
Listing for: oryxsearch.io
Full Time position
Listed on 2026-09-22
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 520000 - 700000 AED Yearly AED 520000.00 700000.00 YEAR
Job Description & How to Apply Below

Principal Evals Engineer - AI & Agentic Systems

About the Organisation

Our client is building one of the region’s most ambitious government AI programmes, developing the AI infrastructure, applications and platforms required to operate AI-native public services at scale.

The systems being built must operate reliably across Arabic and English
, within strict requirements around data sovereignty, privacy, security and regulated infrastructure.

These are production systems supporting consequential workflows and senior decision-makers. Whether an AI system works cannot simply be a matter of opinion.

It has to be measurable.

This role owns how that measurement happens.

Why Join
  • Solve difficult problems in applied AI. Determine whether non-deterministic AI systems actually work across real-world, bilingual and high-stakes environments.
  • Own the quality bar. Define the evidence required before model changes, prompt updates, framework migrations or architectural changes reach production.
  • Small, senior engineering teams. Work alongside experienced AI, product and forward-deployed engineers with significant technical autonomy.
  • Build the evaluation practice. Establish the evaluation architecture, tooling and engineering standards that teams across the programme will use.
  • High-impact environment. Your work will influence AI systems deployed across major public-sector organisations and used at significant scale.
  • Relocation supported. Relocation assistance is available for successful candidates and their families.
The Role

We are hiring a Principal Evals Engineer to own how the organisation determines whether its AI systems actually work.

You will set the direction for evaluation across a portfolio spanning AI assistants, retrieval systems, agentic workflows, voice applications and document intelligence
, while building the infrastructure that turns “it seems better” into measurable evidence.

You own measurement and release evidence
, partnering closely with Principal-level engineers responsible for AI architecture, product engineering, platform development and forward deployment.

Evaluating AI systems is fundamentally different from conventional software testing. The same input can produce different outputs, correctness can be subjective, and failures such as hallucination, poor grounding, behavioural drift and prompt injection cannot be captured through conventional assertions alone.

At Principal level, you will be expected to have already tackled these problems in production environments.

This is primarily an individual contributor role
. Your influence comes from the infrastructure you build, the quality bar you establish and the engineering decisions your evidence enables.

What You Own

Evaluation strategy and architecture.

Define what gets measured, at which layer, using which methodology and how results feed into product and release decisions. Establish the reference architecture engineering teams build against.

The shared evaluation platform.

Build evaluation harnesses, golden-set management, dataset versioning, automated grading, behavioural regression detection and reporting infrastructure.

Grading you can trust.

Own judge-model selection, rubric design and calibration against human labels — including understanding when automated grading cannot be trusted.

Retrieval, agents and multilingual evaluation.

Measure grounding, citation correctness, tool use, multi-step reasoning and failure recovery. Build dedicated Arabic evaluation datasets and ensure judges are properly calibrated rather than assuming English evaluation methods transfer directly.

Evidence behind engineering decisions.

Model swaps, prompt changes, framework migrations and infrastructure decisions should be supported by defensible evaluation evidence before…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary