×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff, RL Systems

Job in Menlo Park, San Mateo County, California, 94029, USA
Listing for: Goaly
Full Time position
Listed on 2026-09-04
Job specializations:
  • Software Development
    Cloud Engineer - Software, DevOps, AI Reliability/ Performance Engineer, Backend Developer
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below

About us

We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.

About

the role

You will build the core platform for agentic reinforcement learning: rollout inference, environment orchestration, distributed training, scheduling, data movement, observability, and recovery. Your goal is to make ambitious RL experiments easy to launch, fast to iterate, efficient to scale, and reliable enough to run for days without constant intervention.

This role sits at the intersection of distributed systems, ML infrastructure, inference, and performance engineering. You will work directly with researchers to find the bottlenecks that matter, then redesign the system so that a local optimization becomes durable leverage for every future run.

What you’ll do
  • Architect and implement end-to-end RL pipelines that coordinate asynchronous rollout generation, environment execution, reward computation, training, evaluation, and checkpoint promotion.

  • Build environment infrastructure for sandboxed and stateful agent workloads, including lifecycle management, isolation, retries, timeouts, replay, checkpoint and restore, and clear distinction between task outcomes and infrastructure failures.

  • Scale high-throughput rollout inference through batching, scheduling, caching, load balancing, disaggregated execution, and efficient model-weight updates.

  • Design resource-management and scheduling systems that place heterogeneous RL workloads efficiently across GPU, CPU, memory, network, and storage constraints.

  • Improve training-inference synchronization, checkpointing, fault recovery, elastic scaling, and long-running job resilience.

  • Profile the full stack and remove bottlenecks in GPU utilization, kernels, communication, serialization, data transfer, storage, and environment throughput.

  • Establish correctness and reproducibility through versioned artifacts, data lineage, idempotent operations, invariant checks, and tests for silent failure modes.

  • Build observability and debugging tools that let researchers answer why a run is slow, unstable, or behaviorally wrong without depending on an infrastructure specialist.

  • Create simple APIs and abstractions that make correct, efficient system use the default while preserving the flexibility needed for fast-moving research.

You may be a good fit if you have
  • A strong record building and operating distributed systems, ML infrastructure, high-performance computing platforms, or performance-critical backend systems.

  • Excellent programming ability in Python plus at least one systems language such as C++, Rust, or Go.

  • Experience reasoning about concurrency, partial failure, back pressure, scheduling, consistency, retries, and observability in production systems.

  • Ability to profile a complex workload, identify the limiting resource, and deliver optimizations that hold under realistic scale and failure conditions.

  • Comfort working across boundaries: research code, model runtimes, infrastructure services, cluster schedulers, and accelerator behavior.

  • Strong ownership and communication, including the ability to turn loosely defined research pain into maintainable platform capabilities.

Strong pluses
  • Experience with distributed RL, large-scale pre-training or post-training, online inference, or asynchronous actor-learner architectures.

  • Familiarity with PyTorch or JAX and systems such as FSDP, Megatron, Deep Speed, Ray, Kubernetes, vLLM, SGLang, TensorRT-LLM, or similar tools.

  • Knowledge of GPU architecture, CUDA or Triton, NCCL/RCCL, RDMA, Infini Band, NVLink, or topology-aware scheduling.

  • Experience building sandboxed execution, workflow engines, actor systems, durable runtimes, or multi-agent orchestration.

  • Meaningful contributions to open-source ML systems or infrastructure projects.

How we work
  • Mission first. We choose work for its…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary