More jobs:
RLHF Manager
Job in
Cheltenham, Gloucestershire, GL50, England, UK
Listed on 2026-09-01
Listing for:
Coaley Peak
Full Time
position Listed on 2026-09-01
Job specializations:
-
IT/Tech
AI Evaluation, AI Engineer (Applied/Software)
Job Description & How to Apply Below
Location Remote (UK) / Cheltenham, UKSalary £55,000 – £75,000 per annum (DOE)
Positions1 position
ReferenceCP-RLF-Published
21 March 2025
Closing
14 December 2026›AI Research & Quality›RLHF Manager About the role Reinforcement Learning from Human Feedback is one of the most important mechanisms for making AI systems behave well in the real world. As RLHF Manager at Coaley Peak, you will design and manage the feedback pipelines, annotation programmes, and evaluation frameworks that keep our models (and our clients' models) aligned with human values, business objectives, and UK regulatory expectations.
This is a dual-facing role. Internally, you will own the RLHF and alignment processes for our proprietary AI engines (Owlpen, Anvil, Flint, Warden), working directly with our data scientists and engineers to define reward models, manage human evaluator programmes, and measure output quality over time. Externally, you will support client AI projects where alignment and quality assurance are a requirement, particularly in regulated sectors such as financial services, healthcare, and government.
We are looking for someone with a rigorous mind, a genuine interest in AI safety and alignment, and the project management capability to run structured annotation and evaluation programmes s role sits within our AI Research & Quality team and reports to the Head of AI.The kind of person we are looking forAt Coaley Peak, the technical work is only half the job.
We are looking for people who are genuinely reliable, who do what they say they will, when they said they would, without needing to be chased. People who are friendly and easy to work with, both with colleagues and with clients. People who can sit in a boardroom and explain a complex AI model in plain English, and then go back to their laptop and write clean, well-documented code.
People who are hungry, hard-working, and take real pride in their output, not because someone is watching, but because that is simply how they operate.
You are meticulous without being slow. You understand that getting AI alignment right requires both technical rigour and human judgement, and you are equally comfortable designing an evaluation rubric, briefing a team of human annotators, and presenting findings to a client's technical steering group. You are intellectually honest, you flag problems before they become failures, and you treat every degradation in model behaviour as a signal worth investigating rather than a number to smooth over.
You care about AI being safe and useful, not just impressive.
- →Support external client AI projects requiring RLHF, alignment review, or model evaluation services, scoping requirements, managing delivery, and reporting outcomes
What we are looking for Essential* →Strong understanding of RLHF, preference learning, and alignment concepts, able to explain and ope rationalise them in a business context* →Experience managing structured data annotation, human evaluation, or labelling programmes* →Proficiency in Python; familiarity with ML frameworks and LLM APIs (OpenAI, Anthropic, Hugging Face, or similar)* →Excellent project management skills: able to run parallel programmes with multiple contributors and tight quality standards* →Clear, precise written and verbal communication, able to document methodology and present findings to both technical and non-technical audiences* →Right to work in the United Kingdom Desirable (not essential)* →Direct experience with RLHF pipelines in production (reward modelling, PPO, DPO, or similar)* →Familiarity with Constitutional AI, RLAIF, or other scalable oversight approaches* →Experience in AI safety research or AI governance* →Background in cognitive science, linguistics, psychology, or philosophy (relevant to preference elicitation and evaluation design)* →Experience working in regulated sectors where AI output…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×