×
Register Here to Apply for Jobs or Post Jobs. X

Machine Learning Infrastructure Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Physical Intelligence
Full Time position
Listed on 2026-09-04
Job specializations:
  • Software Development
    Cloud Engineer - Software, Machine Learning/ ML Engineer, DevOps, AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 180000 - 240000 USD Yearly USD 180000.00 240000.00 YEAR
Job Description & How to Apply Below
  • In this role you will help scale and optimize our training systems and core model code. You’ll own critical infrastructure for large-scale training, from managing GPU/TPU compute and job orchestration to building reusable and efficient JAX training pipelines
  • You’ll work closely with researchers and model engineers to translate ideas into experiments—and those experiments into production training runs
  • This is a hands‑on, high‑leverage role at the intersection of ML, software engineering, and scalable infrastructure
  • The ML Infrastructure team supports and accelerates PI’s core modeling efforts by building the systems that make large‑scale training reliable, reproducible, and fast.
  • The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production‑grade training runs
  • Own training/inference infrastructure:
    Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging
  • Scale distributed training:
    Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction
  • Optimize performance:
    Profile and improve memory usage, device utilization, throughput, and distributed synchronization
  • Enable rapid iteration:
    Build abstractions for launching, monitoring, debugging, and reproducing experiments
  • Manage compute resources:
    Ensure efficient allocation and utilization of cloud-based GPU/TPU compute while controlling cost
  • Partner with researchers:
    Translate research needs into infra capabilities and guide best practices for training at scale
  • Contribute to core training code:
    Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics
Experience & Requirements
  • Experience designing abstractions that balance researcher flexibility with system reliability
  • Hands‑on large‑scale training experience in JAX (preferred), Py Torch
  • Deep ML systems background (e.g., training compilers, runtime optimization, custom kernels)
  • Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS)
  • Experience operating close to hardware (GPU/TPU performance tuning)
  • Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms
  • Familiarity with distributed training, multi‑host setups, data loaders, and evaluation pipelines
  • Bonus Points If You Have Background in robotics, multimodal models, or large‑scale foundation models
  • Ability to debug and optimize performance bottlenecks across the training stack
  • Strong cross‑functional communication and ownership mindset
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary