×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff - Research

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Patronus AI, Inc.
Full Time position
Listed on 2026-08-16
Job specializations:
  • Research/Development
    Research Scientist, AI Evaluation, AI Business & Operations, Data Scientist
Salary/Wage Range or Industry Benchmark: 175000 - 300000 USD Yearly USD 175000.00 300000.00 YEAR
Job Description & How to Apply Below

Patronus AI is a frontier lab developing simulation research and infrastructure to accelerate progress toward human-aligned AGI. We are on a mission to simulate all of the world’s intelligence.

We are the team behind some of the earliest and most influential research in AI evaluation like Finance Bench , Lynx , Simple Safety Tests  , Copyright Catcher , Humanity’s Last Exam , and more. We are formerly AI researchers and engineers from companies like Meta AI, Amazon AGI, and Google. Our customers include foundation model labs and Fortune 500 enterprises like Adobe.

We are backed by top-tier investors like Lightspeed Venture Partners, Notable Capital, Stanford University, Noam Brown, Gokul Rajaram, and more.

Responsibilities

As a Researcher at Patronus AI, you will own and drive foundational research that defines how agentic AI systems are trained, evaluated, and improved. You will work at the intersection of reinforcement learning, simulations, and scalable oversight, building systems that directly influence how frontier models are developed, stress-tested, and deployed.

This is a highly autonomous role. You will tackle open-ended research questions surrounding agent simulations and translate them into rigorous experiments, benchmarks, environments, and production systems. You will work across areas including reward design, tool simulations, agent cognition, behavior analysis, and scalable oversight, helping shape the industry standard for robust, high-quality environments.

Your work will inform how frontier labs design, train, evaluate, and improve the next generation of agents for complex, long-horizon tasks, advancing our path toward safe, human-aligned general intelligence.

In this role, you will:

  • Own ambitious research projects end-to-end
    , from identifying and formulating open-ended problems through experiment design, execution, analysis, and production impact.
  • Advance research in agent simulation, reinforcement learning, and scalable oversight
    , including agent cognition, behavior analysis, reward design, and new training methods.
  • Design state-of-the-art simulation and RL environments for training and evaluating frontier agents, spanning tools and actions, observations and state, trajectories, curricula, and reward systems.
  • Train models and experiment with post-training algorithms
    , including GRPO and SFT. Run ablations with open source models and understand the impact of distillation, COT reasoning, sparse and dense rewards and hyperparameters.
  • Develop methods to understand and improve agent behavior across complex, long-horizon tasks, including reasoning, planning, adaptation, generalization, and reward hacking.
  • Run rigorous experiments and turn findings into measurable outcomes
    , including new techniques, benchmarks, datasets, environments, platform capabilities, and research publications.
  • Build high-quality, reproducible research systems
    , writing production-level code and partnering closely with engineering and product to translate research into real-world systems.
  • Contribute to Patronus AI’s research direction and thought leadership
    , staying at the frontier of the field, collaborating with the research community, and publishing or open sourcing our work.
Qualifications

"The number one qualification to succeed in this machine learning course is gumption” - John Lafferty, CS Professor at Yale

Above all, we look for an eagerness to learn, passion for research, creativity in problem solving and a proactive mindset. You are a great fit if you have a background in the following:

  • An MS or PhD in Computer Science, Machine Learning, Statistics, Mathematics, or a related quantitative field.
  • Experience conducting independent research in reinforcement learning, NLP, agentic systems, evaluation, alignment, or related areas.
  • Demonstrated ability to take open-ended research problems from 0→1 and deliver high-impact outcomes.
  • Strong experimental skills, including experiment design, analysis, and interpretation of results.
  • Experience writing clean, reproducible research code in Python and modern machine learning frameworks.
  • Ability to execute quickly and independently with minimal guidance while maintaining a high bar…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary