×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff - Research Software Engineer - Safety Evaluations Infrastructure

Job in Greater London, London, Greater London, W1B, England, UK
Listing for: United States Digital Space LLC
Full Time position
Listed on 2026-07-30
Job specializations:
  • Software Development
Salary/Wage Range or Industry Benchmark: 90000 - 150000 GBP Yearly GBP 90000.00 150000.00 YEAR
Job Description & How to Apply Below
Location: Greater London

Our Mission

the company is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.

Our Mission

the company is a research lab making intelligence open and accessible for everyone to use, customize, and build on. We build open models that let anyone control their intelligence and help shape the future of AI. Our mission: make intelligence open and accessible to all.

About the Role

As a Research Software Engineer on the Safety team, you will design, build, and own the infrastructure used to run our most sensitive model evaluations — including evaluations in CBRN (chemical, biological, radiological, and nuclear), child safety, and other dangerous-capability domains. These evaluations inform release decisions for our open models, so the systems you build must be secure, isolated, reproducible, and trustworthy under scrutiny.

This is a deeply technical, high-ownership role at the intersection of platform engineering, security, and safety research. You will partner closely with domain experts, legal, and safety researchers to turn their evaluation needs into robust, scalable infrastructure: sandboxed execution environments, controlled data pipelines for sensitive material, access controls, audit logging, and the tooling that lets researchers safely elicit and measure model capabilities in high-consequence areas.

What

You'll Do
  • Design and build secure, sandboxed infrastructure for running sensitive model evaluations, including CBRN and other dangerous-capability domains.
  • Build controlled data pipelines and storage for sensitive evaluation material, applying least-privilege and need-to-know access, role-based access control (RBAC), encryption at rest and in transit, audit logging, and data-minimization safeguards.
  • Partner with safety researchers and domain experts to translate evaluation designs into reliable, reproducible, and scalable systems.
  • Build eval-orchestration tooling and harnesses that let researchers run high-throughput evaluations against models and agents in isolated environments.
  • Develop infrastructure for measuring AI capability uplift in high-consequence domains, and integrate results into the pipelines that inform release decisions.
  • Implement guardrails, monitoring, and compartmentalization so sensitive work stays appropriately siloed, applying least-privilege, need-to-know, and defense-in-depth principles across compute, data, and tooling.
  • Write production-quality Python (and related tooling) for high-throughput data processing and evaluation systems.
  • Improve the reliability, security posture, and developer experience of the safety team's evaluation platform over time.
About You
  • Strong software engineering skills, particularly in Python, with a track record of building reliable, scalable infrastructure or platform systems.
  • Experience building sandboxed, isolated, or otherwise security-sensitive execution environments (e.g., containerization, VM isolation, secure compute) for Trust and Safety teams.
  • Solid grounding in security engineering fundamentals: principle of least privilege, need-to-know access, role-based access control (RBAC), secrets management, encryption, audit logging, compartmentalization, and defense-in-depth design.
  • Experience building data pipelines and handling sensitive or restricted data with appropriate safeguards.
  • Ability to own entire problems end-to-end, including ambiguous, cross-functional ones.
  • Comfort working on sensitive projects that require discretion, integrity, and sound judgment.
  • Thrive in a fast-paced, high-agency startup environment with a bias toward action.
Strong candidates may also have
  • Experience building evaluation, benchmarking, or experimentation infrastructure for ML systems.
  • Experience working with LLMs, agents, or ML training/inference pipelines.
  • Familiarity with dangerous-capability or dual-use domains (CBRN, cyber, etc.) and the information-security considerations they involve.
  • Familiarity with compliance frameworks relevant to sensitive data handling.

We…

Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary