Member of Technical Staff, GPU/ML Systems
Listed on 2026-10-06
-
IT/Tech
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Systems Engineer, Cloud Computing: Infrastructure & Operations
About Sky Pilot
Sky Pilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so Sky Pilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."
Sky Pilot (10k+ Git Hub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.
Therole
Sky Pilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make Sky Pilot the fastest, most cost-efficient place to run demanding AI workloads.
A few points of GPU utilization here can save a team millions in compute and days on every training run.
- Own GPU scheduling, utilization and health
: how Sky Pilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes, with real-time GPU health monitoring and automatic failure recovery. - Build optimizations for training and serving
:
Enable large scale pre-training with node hot-swapping, design storage systems for fast model checkpointing, container migration, inference autoscaling and multi-cluster serving, preemption handling, and sandboxes for training, RL rollouts, and evals. - Make the AI stack run great out of the box
: deepen integrations with vLLM, PyTorch, Slime, and the frameworks teams use for pre-training and high-throughput inference.
- Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure.
- Strongly preferred: Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe).
- You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after.
- Strong Python, and comfort reaching into systems-level and GPU-adjacent details.
- You care about squeezing most from the compute available to you
- Experience operating large-scale training or high-throughput inference in production
- Competitive compensation and equity
- Comprehensive medical, dental, vision coverage for you and your dependents
- The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.
- A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).
- Gourmet lunch & dinner for the team to do their best work
Location: San Mateo, CA. Remote will be considered for exceptional candidates.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).