Member of Technical Staff, AI Infrastructure
Listed on 2026-09-03
-
Software Development
AI Reliability/ Performance Engineer, Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
About us
We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.
About the roleRunning modern AI workloads at scale creates systems problems that rarely fit within a single layer of the stack. A slowdown that appears in a training job may originate in a GPU kernel, collective communication, data movement, container runtime, storage path, scheduler, or interaction between model architecture and hardware topology.
You will identify these bottlenecks and build systems that improve the throughput, efficiency, and robustness of our largest distributed workloads. Your scope will span post-training, agentic reinforcement learning, model training, rollout inference, and the GPU cluster platform beneath them. You will work closely with researchers and systems engineers, develop a quantitative understanding of performance, and turn one-off investigations into durable infrastructure improvements.
This role is a strong fit for an exceptional systems or performance engineer who wants to work on frontier AI infrastructure. Deep ML experience is valuable but not required; we care most about a track record of solving difficult systems problems at scale and the ability to become fluent in new parts of the ML stack quickly.
What you'll doProfile end-to-end AI workloads and identify limiting resources across model code, GPU kernels, memory, collective communication, networking, storage, orchestration, and environment execution.
Build low-latency, high-throughput sampling and inference systems for large language models, including batching, scheduling, caching, load balancing, and efficient weight updates.
Optimize GPU execution through kernel and graph profiling, memory-layout improvements, reduced-precision computation, communication overlap, compilation, and targeted CUDA or Triton work.
Improve distributed training and reinforcement-learning performance across heterogeneous GPU and CPU workloads, variable-length rollouts, complex network topologies, and changing model architectures.
Design quantitative performance and capacity models that predict bottlenecks, explain scaling behavior, guide hardware and topology choices, and prioritize engineering work.
Build scheduling and load-balancing mechanisms that improve accelerator utilization while respecting memory, locality, topology, latency, and fault-domain constraints.
Design fault-tolerant distributed systems that detect failures early, isolate their impact, recover efficiently, and preserve correctness during long-running jobs.
Investigate difficult production issues such as kernel-level stalls, network-latency spikes, collective timeouts, memory fragmentation, stragglers, and performance regressions in containerized environments.
Develop benchmarks, profiling tools, performance dashboards, and regression tests that make system behavior visible and allow improvements to be measured under realistic workloads.
Partner with researchers to understand new models and algorithms, remove infrastructure constraints from the experimental loop, and translate successful optimizations into reusable platform capabilities.
Within your first three months, you have built a quantitative understanding of at least one critical workload, identified its dominant bottlenecks, and shipped a measurable performance or reliability improvement.
Within six to twelve months, you have delivered sustained gains in throughput, accelerator utilization, latency, scaling efficiency, or cost across real training, rollout, or inference workloads.
Performance investigations become faster and more rigorous because the team has better benchmarks, models, profiles, and observability—not just undocumented fixes.
Large distributed jobs run predictably across complex hardware and network topologies,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).