Member of Technical Staff, RL Systems
Listed on 2026-09-04
-
Software Development
Cloud Engineer - Software, DevOps, AI Reliability/ Performance Engineer, Backend Developer
About us
We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.
Aboutthe role
You will build the core platform for agentic reinforcement learning: rollout inference, environment orchestration, distributed training, scheduling, data movement, observability, and recovery. Your goal is to make ambitious RL experiments easy to launch, fast to iterate, efficient to scale, and reliable enough to run for days without constant intervention.
This role sits at the intersection of distributed systems, ML infrastructure, inference, and performance engineering. You will work directly with researchers to find the bottlenecks that matter, then redesign the system so that a local optimization becomes durable leverage for every future run.
What you’ll doArchitect and implement end-to-end RL pipelines that coordinate asynchronous rollout generation, environment execution, reward computation, training, evaluation, and checkpoint promotion.
Build environment infrastructure for sandboxed and stateful agent workloads, including lifecycle management, isolation, retries, timeouts, replay, checkpoint and restore, and clear distinction between task outcomes and infrastructure failures.
Scale high-throughput rollout inference through batching, scheduling, caching, load balancing, disaggregated execution, and efficient model-weight updates.
Design resource-management and scheduling systems that place heterogeneous RL workloads efficiently across GPU, CPU, memory, network, and storage constraints.
Improve training-inference synchronization, checkpointing, fault recovery, elastic scaling, and long-running job resilience.
Profile the full stack and remove bottlenecks in GPU utilization, kernels, communication, serialization, data transfer, storage, and environment throughput.
Establish correctness and reproducibility through versioned artifacts, data lineage, idempotent operations, invariant checks, and tests for silent failure modes.
Build observability and debugging tools that let researchers answer why a run is slow, unstable, or behaviorally wrong without depending on an infrastructure specialist.
Create simple APIs and abstractions that make correct, efficient system use the default while preserving the flexibility needed for fast-moving research.
A strong record building and operating distributed systems, ML infrastructure, high-performance computing platforms, or performance-critical backend systems.
Excellent programming ability in Python plus at least one systems language such as C++, Rust, or Go.
Experience reasoning about concurrency, partial failure, back pressure, scheduling, consistency, retries, and observability in production systems.
Ability to profile a complex workload, identify the limiting resource, and deliver optimizations that hold under realistic scale and failure conditions.
Comfort working across boundaries: research code, model runtimes, infrastructure services, cluster schedulers, and accelerator behavior.
Strong ownership and communication, including the ability to turn loosely defined research pain into maintainable platform capabilities.
Experience with distributed RL, large-scale pre-training or post-training, online inference, or asynchronous actor-learner architectures.
Familiarity with PyTorch or JAX and systems such as FSDP, Megatron, Deep Speed, Ray, Kubernetes, vLLM, SGLang, TensorRT-LLM, or similar tools.
Knowledge of GPU architecture, CUDA or Triton, NCCL/RCCL, RDMA, Infini Band, NVLink, or topology-aware scheduling.
Experience building sandboxed execution, workflow engines, actor systems, durable runtimes, or multi-agent orchestration.
Meaningful contributions to open-source ML systems or infrastructure projects.
Mission first. We choose work for its…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).