Research Engineer, Code RL (Reinforcement Learning
Listed on 2026-08-22
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Software Engineer
Location: New York
About Anthropic
Anthropic’s mission is to create reliable, interpretable, and steerable AI systems that are safe and beneficial for users and society.
About the RL TeamsThe Reinforcement Learning teams are critical to advancing our AI systems, contributing to all Claude models and impacting autonomy and coding capabilities. Core work includes:
- Developing systems that enable models to use computers effectively
- Advancing code generation through reinforcement learning
- Pioneering fundamental RL research for large language models
- Building scalable RL infrastructure and training methodologies
- Enhancing model reasoning capabilities
The Role
We are hiring a Research Engineer for the Code RL team. You will design RL environments and coding tasks, build reward signals and verifiers that capture the essence of "good code", run training experiments on frontier models, diagnose model behavior, and improve speed and reliability of training pipelines. The role combines research and engineering across multiple focus areas, including agentic coding behaviors, long‑horizon autonomous engineering, and high‑performance code for accelerators.
YouMay Be a Good Fit If You
- Have strong software‑engineering skills and deep Python expertise, including async and concurrent programming.
- Own systems end to end and debug across the stack.
- Balance research exploration with engineering implementation, and rigorously design experiments and interpret results.
- Care about code quality, testing, and performance.
- Are passionate about AI’s impact and committed to building safe and beneficial systems.
- Experience with reinforcement learning, RLHF, post‑training, or LLM fine tuning.
- Built coding agents, code‑execution sandboxes, evaluation harnesses, verifiers, or developer tooling.
- Background in program analysis, testing, verification, compilers, or formal methods.
- Experience with PyTorch and large‑scale distributed training; performance profiling and ML system optimization.
- CUDA / GPU or TPU kernel experience and accelerator‑performance intuition.
- Experience with virtualization and sandboxed code execution environments.
- Research Engineer, Performance RL — teach Claude to write correct, fast code for accelerators.
- Research Engineer, Universes — long‑horizon, ultra‑realistic agentic training environments.
- Research Engineer, Cybersecurity RL — RL for security‑relevant coding capabilities.
$500,000—$850,000 USD
Logistics- Minimum education:
Bachelor’s degree or equivalent combination of education, training, and experience. - Required field of study:
Relevant coursework, training, or professional experience. - Minimum years of experience:
Varies by internal job level. - Location‑based hybrid policy:
Staff are expected to be in an office at least 25% of the time. - Visa sponsorship:
We sponsor visas where possible and will make every reasonable effort to obtain a visa for an offer recipient.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).