×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff, GPU​/ML Systems

Job in San Mateo, San Mateo County, California, 94409, USA
Listing for: SkyPilot
Full Time position
Listed on 2026-10-06
Job specializations:
  • IT/Tech
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 180000 - 280000 USD Yearly USD 180000.00 280000.00 YEAR
Job Description & How to Apply Below
Position: Member of Technical Staff, GPU / ML Systems

About Sky Pilot

Sky Pilot accelerates the world's most ambitious AI teams. Every hour they spend fighting infrastructure is an hour the frontier doesn't move — so Sky Pilot turns fragmented compute across clusters into one optimized, highly available and easy-to-use pool: a single "AI supercomputer."

Sky Pilot (10k+ Git Hub stars, 14M+ downloads) is deployed at 100s of companies — from Fortune 500s to top AI-natives like Abridge, Applied Compute, Mistral, Unconventional AI, H Company, and Nubank — with usage growing exponentially. Born in the UC Berkeley lab behind Spark and Databricks, our growing team includes top-tier talent from Databricks, Google, Berkeley, MIT, CMU, and Cornell.

The

role

Sky Pilot exists because GPUs are scarce, expensive, and scattered — and the workloads that need them (pre-training, post-training, RL, high-throughput inference) push hardware to its limits. We're looking for an engineer to own the GPU and ML-systems layer that frontier AI teams run on: accelerator scheduling and utilization, the serving path, and the integrations that make Sky Pilot the fastest, most cost-efficient place to run demanding AI workloads.

A few points of GPU utilization here can save a team millions in compute and days on every training run.

What you'll do
  • Own GPU scheduling, utilization and health
    : how Sky Pilot discovers, packs, and binpacks accelerator capacity across clouds and Kubernetes, with real-time GPU health monitoring and automatic failure recovery.
  • Build optimizations for training and serving
    :
    Enable large scale pre-training with node hot-swapping, design storage systems for fast model checkpointing, container migration, inference autoscaling and multi-cluster serving, preemption handling, and sandboxes for training, RL rollouts, and evals.
  • Make the AI stack run great out of the box
    : deepen integrations with vLLM, PyTorch, Slime, and the frameworks teams use for pre-training and high-throughput inference.
What we're looking for
  • Hands-on experience with GPU or accelerator systems and with ML training or inference infrastructure.
  • Strongly preferred: Familiarity with the modern ML ecosystem (e.g. vLLM, PyTorch, CUDA, verl/slime) and workload-orchestration frameworks (e.g. Kueue, KAI, KServe).
  • You've done real ML-systems performance work - tell us about a bottleneck you hunted down (a stalled data pipeline, GPUs idling on a scheduling gap, communication you overlapped with compute) and what you measured before and after.
  • Strong Python, and comfort reaching into systems-level and GPU-adjacent details.
  • You care about squeezing most from the compute available to you
  • Experience operating large-scale training or high-throughput inference in production
What we offer
  • Competitive compensation and equity
  • Comprehensive medical, dental, vision coverage for you and your dependents
  • The chance to work with some of the best minds in cloud, distributed, and AI systems — with significant autonomy and ownership.
  • A front-row seat at the latest open-source infra startup from Berkeley (lineage: Databricks, Anyscale).
  • Gourmet lunch & dinner for the team to do their best work

Location: San Mateo, CA. Remote will be considered for exceptional candidates.

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary