More jobs:
Research Scientist – RL Post-Training Agents
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-10-05
Listing for:
Rnb Consultancy
Apprenticeship/Internship
position Listed on 2026-10-05
Job specializations:
-
Research/Development
Research Scientist, AI Business & Operations
Job Description & How to Apply Below
Want your RL research to land in agents that run for days in the real world, not in a paper appendix?
A brand-new AI lab in San Francisco is building autonomous agents that pursue complex goals over very long horizons. The founding team comes from leading frontier AI labs, autonomous-driving and robotics AI, big-tech research and a top quant firm. This is a research seat that builds real systems: you own ambitious bets from the first hypothesis and dataset all the way to a deployed capability.
They’re hiring 4 research scientists.
What you’ll own- Using RL, and whatever else works, to post-train LLM-based and multimodal agents
- Long-horizon capabilities: planning, memory, error recovery, proactivity, self-improvement, and knowing when to escape to a human
- Building the environments where agents use computers and tools, plus the data pipelines, benchmarks and evals around them
- Ways for agents to grasp what a user wants and stay true to it over long runs
- Careful experiment design, and taking what works all the way into production
- A real say in the research agenda from day one
- Exceptional ML research and engineering skills
- Depth in at least one of these: RL, LLM post-training, reasoning, agents, computer use, long-horizon systems, memory and context, evals, or how humans and agents work together
- Good judgment on what to test and when to stop. You build complete systems and move easily between ideas, large-scale experiments and production.
- You’ve led something significant: a model, an agent system, a benchmark, a paper, an open-source project or a big research bet
- Roughly 3–6 years in; frontier-lab experience valued, exceptional outliers and senior leads welcome
- High agency and comfort with uncertain directions
- Hands-on RL post-training or computer-use work
- A strong publication, open-source or benchmark record
- Experience building environments and eval harnesses
- A spike: olympiad (IOI/IMO), quant, or world-class competitive achievement
- $250k–$500k base + 1–5% equity
- Your own research bets, end to end, at a lab where results ship
- Visa sponsorship available; if you’re outside the US, expect to go via an O-1
- Full-time, in person in San Francisco, 9-9-6
- Process: informal talk with a founder → technical deep-dive on your research → paid 2–3 day work trial in person → offer
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×