Software Engineer, Compute Infrastructure
Listed on 2026-09-10
-
Software Development
Unix/Linux, DevOps, Software Engineer, Cloud Engineer - Software
Role Overview
The role belongs to a Software Engineer focused on Compute Infrastructure. Candidates joining this position build the platform that converts massive compute capacity into a dependable engine for frontier AI. This position designs, provisions, schedules, operates, and optimizes the systems that connect accelerators, CPUs, networks, storage, data centers, orchestration software, agent infrastructure, developer tools, and observability into one coherent experience for researchers and product teams.
Scopeof Impact
This role operates across the entire Compute Infrastructure stack. Responsibilities include capacity planning and cluster lifecycle, bare-metal automation, distributed systems, Kubernetes and scheduling, deep system optimization, high-performance networking, storage, fleet health, reliability, workload profiling, benchmarking, and the developer experience. At this scale, improvements in communication, scheduling, hardware efficiency, or debugging workflows can compound into meaningful research velocity. The company hires across Compute Infrastructure rather than for a single narrow team, using this opening to match strong engineers to problems where they can generate the most leverage.
PotentialWork Areas
Candidates may work close to hardware or close to users, on CaaS and agent infrastructure, or on the control planes and data planes in between. You might help bring new supercomputing capacity online, optimize training workloads from profiler traces and benchmarks, improve NCCL and collective communication behavior, reason about GPUs, NICs, topology, firmware, thermals, and failure modes, or design abstractions that make heterogeneous clusters feel like one coherent platform.
ExpectedContributions
You will build and deeply optimize reliable system software for large-scale compute systems that run some of the world's most demanding AI workloads. You will design and operate infrastructure across accelerators, CPUs, NICs, switches, networking protocols, storage, data centers, cluster orchestration, scheduling, and fleet health. Profiling, benchmarking, and optimization of training workloads across compute, memory, storage, networking, NCCL and collective communication, and cluster scheduling bottlenecks will form part of your daily work.
You will create hardware‑aware automation that makes provisioning, firmware and driver upgrades, incident response, and day‑to‑day operations faster and less error‑prone. Building CaaS, agent infrastructure, profiling, observability, benchmarking, and platform tools that help researchers, product engineers, and operators launch, debug, and optimize workloads with less friction will be central. You will turn operational lessons into better systems, stronger abstractions, and clearer ownership boundaries across teams.
Collaboration across research, engineering, security, networking, hardware, and data center teams will help make compute capacity more capable and easier to use.
You have built or operated distributed systems, infrastructure platforms, high‑performance computing environments, large‑scale networking systems, Kubernetes clusters, developer tools, or production systems with demanding reliability requirements. You enjoy working across layers of the stack and are comfortable moving between software, hardware, networking, systems performance, reliability, and user needs. You care about making complex infrastructure understandable, observable, and usable for the people depending on it.
You can diagnose hard problems under real operational pressure while still investing in long‑term engineering quality. You like building leverage for others, whether through APIs, automation, debugging tools, CaaS and agent…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).