Software Engineer, Inference
Listed on 2026-08-07
-
Software Development
DevOps
You’ll own how Luma’s models get served — integrating new architectures into the inference engine, scaling deployments across thousands of machines, and keeping expensive GPU fleets busy while meeting internal SLOs.
This is large-scale inference systems work: scheduling, fleet management, deployment pipelines, and reliability across clusters and hardware providers. It fits a strong systems engineer comfortable with model serving and Kubernetes you want pure modeling rather than the systems that run models, this is firmly the systems side.
What You’ll Own- Ship new model architectures by integrating them into the inference engine.
- Collaborate across research, engineering, and infrastructure to optimize model efficiency and deployments.
- Build internal tooling to measure, profile, and track the lifetime of inference jobs and workflows.
- Automate, test, and maintain inference services for maximum uptime and reliability.
- Manage and optimize inference workloads across clusters and hardware providers, and scale deployments across thousands of machines.
- Build scheduling systems that use expensive GPU resources optimally while meeting SLOs, and maintain CI/CD for model checkpoints and SDKs.
One way the first 90 could unfold.
- Days 1–30 — Immerse & Diagnose: Learn the inference stack, the fleets, and where reliability or utilization break.
- Days 30–60 — Ship & Validate: Integrate a model or ship tooling/scheduling that improves uptime or GPU utilization.
- Days 60–90 — Scale & Systemize: Harden deployment pipelines and scheduling across clusters and providers.
- Strong Python and system-architecture skills.
- Experience deploying models with PyTorch, Hugging Face, vLLM, SGLang, TensorRT-LLM, or similar.
- Experience with queues, scheduling, traffic control, and fleet management at scale.
- Experience with Linux, Docker, and Kubernetes, and with orchestration, deployment, and scheduling.
- Familiarity with Redis and S3-compatible storage.
- Modern networking stacks including RDMA (RoCE, Infini Band, NVLink).
- High-performance large-scale ML systems (100+ GPUs).
- CUDA, and FFmpeg or multimedia processing.
About Luma:
Luma’s mission is to build unified general intelligence that can generate, understand, and operate in the physical world. We believe multimodality is critical for intelligence — the next step beyond language models comes from vision. Luma is an equal opportunity employer.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).