×
Register Here to Apply for Jobs or Post Jobs. X

Founding Member of Technical Staff

Job in Greater London, London, Greater London, W1B, England, UK
Listing for: Sampura
Full Time position
Listed on 2026-09-10
Job specializations:
  • Research/Development
    AI Evaluation, Research Scientist, Data Scientist
Salary/Wage Range or Industry Benchmark: 100000 - 290000 GBP Yearly GBP 100000.00 290000.00 YEAR
Job Description & How to Apply Below
Location: Greater London

About

Sampura Research is an AI safety research non-profit focused on scalable oversight for frontier models through improved human/AI “judges”.

Judges (LLM, human, or a hybrid) are essential for aligning highly capable models through their use in evaluation and training. Limitations in today's judges can lead to undesirable behavior such as reward hacking in training, or failure to spot misalignment in evaluation. As models become more capable and their data more complex, often beyond what the systems overseeing them can understand, we expect these limitations to grow in severity and to become harder to notice.

We are designing novel judge protocols which leverage the strengths of both humans and models to outperform either alone, and building high-quality evaluations to explore judge performance across a diverse set of safety-related tasks. We believe that our focused approach will rapidly advance the science of judges by providing a much-needed systematic measurement of which protocols actually work, and producing state-of-the‑art judges which can be used in real world training and evaluation pipelines.

Backed by an $11M grant from Coefficient Giving ($7M for the first year, and another $4M pledged),
Sampura Research was founded in 2026 by members of Deep Mind’s alignment and human data teams.

Please see our announcement for more information.

Role Overview

As a founding Member of Technical Staff, you will help shape the research and engineering direction alongside the founding team, and own end-to-end research and engineering problems within our agenda:

Evaluation
:
Build diverse and high-quality evaluations for judge protocols by pulling from existing datasets/literature and creating novel datasets from scratch.

Methods
:
Research and implement strategies for improving judge performance through hybridization, such as learning-to-defer, human assistance, or other mechanisms.

Infrastructure
:
Build the tools required to enable and scale our research agenda, including large-scale inference/evaluation pipelines, a human rating platform, and a robust public leader board.

We've budgeted roughly $1M per quarter for compute and human data, across a team of fewer than ten, to ensure that we can commit to ambitious goals in all of these areas. You'll scope ambiguous problems independently, communicate progress clearly, and synthesize findings into papers, blog posts, and benchmark releases.

Staff hires may also lead research and mentorship programs through paid fellowships or volunteer research focused on evaluation or methods work.

If our work and mission resonate with you, please consider applying even if you don’t tick every box. We strongly believe in investing in excellent people regardless of prior or formal experience.

You may be a good fit if you have

  • A Bachelor's degree or higher in Computer Science, Mathematics, or a related field or equivalent experience.

  • Proficiency coding in Python or similar languages, as well as machine learning and analysis tooling (e.g., JAX, PyTorch, Pandas).

  • Experience working with large language models, human-in-the-loop data annotation systems, or AI evaluation frameworks.

  • A willingness to embrace AI tools intelligently, applying scrutiny and human review where necessary.

  • A track record of exploring and resolving open-ended research questions with rigor, in academic or industry labs, fellowships, or volunteer research programs.

Outstanding candidates will have

  • Demonstrated ownership across multiple parts of the AI research lifecycle (experiment design and execution, infrastructure, data collection and processing, inference).

  • Expertise in a highly relevant domain such as model confidence/calibration, interpretability, safety/alignment evaluation, human-computer interaction,…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary