More jobs:
AI Evaluation Engineer
Job in
London, Greater London, W1B, England, UK
Listed on 2026-07-29
Listing for:
Vallum Associates
Full Time
position Listed on 2026-07-29
Job specializations:
-
IT/Tech
AI Evaluation, AI Engineer (Applied/Software)
Job Description & How to Apply Below
AI Evaluation Engineer
Work Location :
London
Type work :
Contract
Mode of work :
Hybrid
Responsibilities:
* Define and implement end-to-end evaluation strategy for generative AI conversational systems (LLMs, RAG pipelines, agents).
* Establish evaluation metrics
* Test Dataset & Benchmarking
* Benchmark models (GPT variants, open-source LLMs) and prompt strategies
* Implement automated evaluation pipelines
* Conduct human-in-the-loop evaluation for qualitative validation.
* Prompt & Response Quality Optimization
* RAG & Knowledge Grounding Validation
* Safety, Risk & Compliance Testing
Your Profile
* Strong understanding of LLMs, RAG architecture, prompt engineering
* Python – data analysis, evaluation pipelines
* Prompt evaluation tools (Prompt Tools, Deep Eval, etc.)
Experience with :
- Evaluation Framework Design
- Establish Evaluation metrics & NLP quality assessment
- Test Dataset and benchmarking
- LLM output evaluation and scoring
- Safety, risk and compliance testing
- Python, data analysis, experimentation frameworks
- Tooling and automation
Familiarity with:
- Responsible AI principles
- Customer support / retail conversational flows
- Analytical mindset with ability to translate model behaviour into actionable improvements
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×