Research Engineer, Infrastructure, RL Systems
Listed on 2026-07-22
-
Software Development
Machine Learning/ ML Engineer, DevOps, Cloud Engineer - Software, Software Engineer
About the Role
We’re looking for an infrastructure research engineer to design and build the core systems that enable scalable, efficient training of large models through reinforcement learning.
This role sits at the intersection of research and large‑scale systems engineering: a builder who understands both the algorithms behind RL and the realities of distributed training and inference at scale.
You’ll wear many hats, from optimizing rollout and reward pipelines to enhancing reliability, observability, and orchestration, collaborating closely with researchers and infra teams to make reinforcement learning stable, fast, and production‑ready.
What You’ll Do- Design, build, and optimize the infrastructure that powers large‑scale reinforcement learning and post‑training workloads.
- Improve the reliability and scalability of RL training pipelines, distributed RL workloads, and training throughput.
- Develop shared monitoring and observability tools to ensure high uptime, debuggability, and reproducibility for RL systems.
- Collaborate with researchers to translate algorithmic ideas into production‑grade training pipelines.
- Build evaluation and benchmarking infrastructure that measures model progress on helpfulness, safety, and factuality.
- Publish and share learnings through internal documentation, open‑source libraries, or technical reports that advance the field of scalable AI infrastructure.
Minimum qualifications
- Bachelor’s degree or equivalent experience in computer science, electrical engineering, statistics, machine learning, physics, robotics, or similar.
- Strong engineering skills, ability to contribute performant, maintainable code and debug in complex codebases.
- Understanding of deep learning frameworks (e.g., PyTorch, JAX) and their underlying system architectures.
- Thrives in a highly collaborative environment involving many cross‑functional partners and subject matter experts.
- A bias for action with a mindset to take initiative to work across different stacks and teams where you spot opportunities to ship.
Preferred qualifications
- Experience training or supporting large‑scale language models with tens of billions of parameters or more.
- Experience working with reinforcement learning workloads (e.g., PPO, DPO, RLHF, or reward modeling).
- Background in high‑performance or reliability engineering – distributed training frameworks and cluster orchestration (Kubernetes, Slurm).
- Familiarity with monitoring and observability tools (Prometheus, Grafana, Open Telemetry).
- Contributions to large‑scale ML research or infrastructure, open‑source frameworks, or internal performance optimization efforts.
This role is based in San Francisco, California.
CompensationDepending on background, skills and experience, the expected annual salary range for this position is $350,000 - $475,000 USD.
Visa SponsorshipWe sponsor visas.
BenefitsThinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
As set forth in Thinking Machines’ Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law. Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).