×
Register Here to Apply for Jobs or Post Jobs. X

Research Lead - Pre-training Safety FAR.AI · Remote · US · AI Research –ago

Remote / Online - Candidates ideally in
Berkeley, Alameda County, California, 94709, USA
Listing for: Aimlroles
Full Time, Apprenticeship/Internship, Remote/Work from Home position
Listed on 2026-10-07
Job specializations:
  • IT/Tech
    AI Business & Operations, AI Engineer (Applied/Software), Data Scientist
  • Research/Development
    AI Business & Operations, Data Scientist
Salary/Wage Range or Industry Benchmark: 290000 - 450000 USD Yearly USD 290000.00 450000.00 YEAR
Job Description & How to Apply Below
About Us

FAR.AI is a non-profit AI research institute working to ensure advanced AI is safe and beneficial for everyone. Our mission is to facilitate breakthrough AI safety research, advance global understanding of AI risks and solutions, and foster a coordinated global response.
Since our founding in July 2022, we've grown to 50+ staff, published 40+ academic papers, and convened leading AI safety events. Our work is recognized globally, with publications at premier venues such as NeurIPS, ICML including a Best Paper Honorable Mention in 2026, and ICLR, and features in the Financial Times, Nature News, Wired Magazine and MIT Technology Review. We conduct pre-deployment testing on behalf of frontier developers such as OpenAI, independent evaluations for governments including the EU AI Office, and publish the AI Security Leader board based on our red-teaming expertise.

We help steer and grow the AI safety field through developing research roadmaps with renowned researchers such as Yoshua Bengio; running FAR.Labs, an AI safety-focused co-working space in Berkeley housing 40+ members; and supporting the community through targeted grants to technical researchers.

FAR.AI is hiring a Research Lead to develop and lead our work on pre-training safety
, shaping models’ capabilities and internal representations at their source, rather than trying to fix them after the fact.

Our initial focus is capability control: removing harmful capabilities while preserving benign ones. We see this as a promising way to prevent misuse of open-weight models in areas such as CBRN and cyber by removing offensive capabilities, and reducing loss-of-control risks by removing knowledge of oversight mechanisms. We will validate approaches like pre-training data filtering at scale, drive adoption of successful methods, and explore techniques such as gradient routing and unlearning.

We are scaling methods like Deep Ignorance by over an order of magnitude (>100B parameter models with >1T tokens). You will direct this work, partner with our red team to stress-test the resulting models, and analyze how well the methods scale to frontier systems.

Our research directions include:

  • Improved data filtering methods, such as using data attribution (e.g. influence-based selection) or more sophisticated classifiers

  • Using methods like gradient routing to isolate dual-use capabilities in components of the model (e.g. specific MoE experts)

  • Training to actively remove harmful capabilities, such as interleaving next-token prediction with unlearning, as opposed to simply filtering data

  • Adding synthetic data to pre-training or mid-training to shape the representations and behavior of the model

You’ll build and lead the team, set its research direction, mentor Members of Technical Staff to scale your vision, and remain hands‑on enough to write code and run experiments yourself. This role offers high autonomy in an impact-driven environment, pursuing empirically grounded, scalable ML safety research.

About FAR.Research

We explore promising research directions in AI safety and scale up only those showing a high potential for impact. When an approach proves effective, we develop it into a minimum viable demonstration and work with AI developers and governments to support real-world adoption.

Our recent and ongoing research includes:

Adversarial Robustness: working to rigorously solve security problems through building a science of security and robustness for AI, from demonstrating superhuman systems can be vulnerable, to scaling laws for robustness and jail breaking constitutional classifiers.

Mechanistic Interpretability: finding issues with Sparse Autoencoders, probing deception using Among Us, understanding learned planning in Soko Ban, and…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary