×
Register Here to Apply for Jobs or Post Jobs. X

Research Scientist ​/ Engineer – Training Infrastructure

Job in Redwood City, San Mateo County, California, 94061, USA
Listing for: Luma AI
Apprenticeship/Internship position
Listed on 2026-08-05
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer
Job Description & How to Apply Below

Distributed Systems Engineer

You'll build the distributed systems that train Luma's large-scale multimodal models across thousands of GPUs, so researchers can focus on innovation on top of reliable, efficient, scalable infrastructure.

This is hard PyTorch, CUDA, and distributed-systems work — advanced parallelism, training stability, and utilization across massive clusters. It fits an engineer who's solved real problems training foundation models  you haven't worked at the level of FSDP and multi-node training, this is the wrong depth.

What You'll Own
  • Design, implement, and optimize efficient distributed training systems for models across thousands of GPUs.
  • Research and implement advanced parallelization (FSDP, Tensor Parallel, Pipeline Parallel, Expert Parallel).
  • Build monitoring, visualization, and debugging tools for large-scale training runs.
  • Optimize training stability, convergence, and resource utilization across massive clusters.
First 90 Days

One way the first 90 could unfold.

  • Days 1–30 — Immerse & Diagnose: Learn the current training stack and where stability and utilization hurt at scale.
  • Days 30–60 — Ship & Validate: Land a parallelization or stability improvement that measurably helps a real training run.
  • Days 60–90 — Scale & Systemize: Build the monitoring and tooling that keeps large runs reliable and efficient.
What You Bring
  • Extensive distributed PyTorch training and parallelisms in foundation-model training.
  • Deep understanding of GPU clusters, networking, and storage systems.
  • Familiarity with communication libraries (NCCL, MPI) and distributed-system optimization.
Nice to Have
  • Strong Linux systems administration and scripting.
  • Experience managing training runs across 100+ GPUs.
  • Experience with containerization, orchestration, and cloud infrastructure.

About Luma:
Luma's mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary