Research Engineer, Infrastructure
Listed on 2026-08-13
-
Software Development
Machine Learning/ ML Engineer, Software Engineer, AI Reliability/ Performance Engineer, AI Engineer (Applied/Software)
Fleet studies how environments produce intelligence. We believe intelligence is an emergent property of environmental pressures: the environment determines what capabilities develop, what behaviors survive, and what "good" looks like. We work with frontier labs on post-training across modalities, building benchmarks that expose where frontier models break, training recipes that close those gaps, and scalable oversight for long-horizon agents. Backed by Sequoia Capital, Menlo Ventures, BCV, and SV Angel.
Aboutthe Role
Fleet runs a large GPU cluster. As a Research Engineer, you own the infrastructure our research runs on: the training stack, the inference stack, the research codebase that ties them together, the cluster itself, and the kernels at the bottom. Your job is to keep all of it fast and reliable as the research moves.
You work close to the researchers and fix what is actually slow, where the fix belongs. Sometimes that is a Triton kernel. Sometimes it is a parallelism config change. Sometimes it is deleting a library nobody is using anymore.
The research codebase is under active development and most of the bottlenecks are not obvious from outside it. We want someone who can profile a stalled collective, write a fused kernel, and refactor research code without breaking the science. This is hire #1 for the role.
What You'll Do- Own the cluster. Capacity, failover, fault tolerance, and on-call when training stalls.
- Drive utilization. Profile training runs, find the bottleneck (NCCL, KV cache, data loader, allgather), and fix it. Write kernels when the profiler points there.
- Maintain training and inference. Distributed training (FSDP, TP, PP), rollout generation, async training, and the inference stack the training loop depends on.
- Maintain the research codebase. Keep it fast and reliable as researchers extend it. Refactor without breaking the science.
- Small, technical team.
- Speed with rigor.
- Truth over comfort.
- On-site in SF.
- Distributed training depth. Hands-on experience with multi-node training (tens of GPUs at minimum, ideally hundreds) and the major parallelism strategies (DP, TP, PP, FSDP).
- Performance and kernels. You can profile a slow training run and land the fix. You write CUDA and Triton, and have shipped a fused op or attention variant in a real training run.
- Inference systems. You have built or maintained a high-throughput inference stack (vLLM, SGLang). KV cache, continuous batching, speculative decoding, and quantization are tools you have used in production.
- Research-codebase taste. You have walked into a sprawling, half-documented research repo and made it measurably faster or more reliable without breaking the science. RL post-training methods (PPO, GRPO, RLHF) and their failure modes are familiar enough that you can triage a training failure from a log line.
- Contributions to PyTorch, vLLM, SGLang, Megatron, Deep Speed, FSDP, NCCL, flash-attention, or other open-source ML systems.
- Experience operating GPU fleets at scale.
- Experience with RoCE / Infini Band, GPUDirect, or similar high-performance networking.
- Prior work on RL rollout systems, async training, or off-policy training infra.
Highly competitive salary and equity. On-site in San Francisco.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).