×
Register Here to Apply for Jobs or Post Jobs. X

Research Scientist, Reinforcement Learning

Job in Palo Alto, Santa Clara County, California, 94306, USA
Listing for: Pramaana Labs
Full Time position
Listed on 2026-09-28
Job specializations:
  • Research/Development
    AI Evaluation, AI Business & Operations
  • IT/Tech
    AI Evaluation, AI Business & Operations, Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 150000 - 230000 USD Yearly USD 150000.00 230000.00 YEAR
Job Description & How to Apply Below

Company

Pramaana Labs is a frontier AI lab building the verification layer for AI. Founded in 2025 and headquartered in Palo Alto, we turn complex human knowledge, including tax codes, legal rules, clinical guidelines and security protocols, into a formal representation, so every AI answer can be traced, challenged, and proved.

Pramaana's architecture pairs foundation models trained to formalize and reason with a symbolic world model encoded in Lean
4. Pramaana was founded by a team out of Google, Deep Mind, and Glean, combining frontier AI researchers and formal methods experts.

The role

We're training foundation models to natively interact with symbolic world models. As an RL Post-Training Researcher, you'll take foundation models and scale their reasoning capabilities: applying RLVR to new domains using verified rewards from the Lean kernel, pushing the frontier of autoformalization and proving, and innovating on RL algorithms, data, and evals. Your work will also define how our models leverage test-time compute to solve long-horizon logical tasks.

What you'll actually work on
  • Scale RLVR to new, real-world domains using the Lean kernel as a deterministic reward signal.

  • Design and implement novel RL algorithms and test-time compute optimizations tailored for formal environments and proof search.

  • Push the frontier of autoformalization, training models to map highly technical natural language into strict formal specifications.

  • Curate data and build evals that tightly correlate with verifiable downstream reasoning capabilities.

  • Own the whole loop: rollout sampling, reward design, the update, the eval.

You Will Thrive in This Role If You Have
  • Deep, hands-on research experience in reinforcement learning applied to reasoning models at scale.

  • Experience working with RLVR, test-time RL, or exact deterministic reward signals.

  • Strong algorithmic and systems intuition — comfortable writing custom RL loops, managing data pipelines, and building robust evals from scratch.

  • High autonomy: the ability to take a fuzzy problem area and independently drive it to state-of-the-art results without day-to-day direction.

Especially Strong Candidates May Also Have
  • A background in formal verification (Lean, Coq, or Isabelle)

  • A strong publication record in top-tier venues (NeurIPS, ICML, ICLR).

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary