Artificial Intelligence Engineer
Listed on 2026-09-27
-
Software Development
Machine Learning/ ML Engineer
Senior Pretraining Engineer – Video & Multimodal AI
Location:Bay Area, CA | Onsite
Focus:Large-scale pre-training, video, multimodal models, diffusion, flow matching
I'm working with an early-stage robotics company building a deeply integrated AI and robotics stack from first principles.
Their long-term goal is ambitious: create highly autonomous factories capable of manufacturing physical goods with dramatically less human manual labour. That means solving problems across robot learning, perception, simulation, rendering, GPU performance and large-scale multimodal model training.
They’re now hiring a
Senior Pretraining Engineer
to take ownership of large-scale video and multimodal pre-training.
This is not a role for someone who has only fine-tuned existing models or operated clean, established training pipelines.
You’ll be expected to understand what happens when
big models, big datasets and big compute collide
, including the failure modes that only become visible once training runs become genuinely expensive.
You’ll work across model architecture, training infrastructure, data and experimentation to build large-scale vision and multimodal systems that can ultimately contribute to robotic intelligence in complex physical environments.
- Own large-scale video and multimodal pre-training runs
- Train models across substantial distributed GPU infrastructure
- Develop and improve generative approaches including
flow matching and diffusion - Diagnose instability, convergence issues and large-scale training failures
- Improve training efficiency, reliability and experiment velocity
- Build safeguards based on lessons from failed or expensive training runs
- Work closely with robotics, simulation, perception and systems engineers
- Push research ideas into working systems rather than isolated experiments
Strong candidates will have experience with:
- Large-scale
video, vision or multimodal pre-training - Distributed training across substantial GPU compute
- Training large models on large datasets
- Debugging difficult training failures at scale
- Understanding the interaction between models, data and compute
- Flow matching, diffusion models or adjacent generative-model approaches
- Strong software engineering and training-systems fundamentals
- Designing experiments where compute cost makes mistakes consequential
Just as importantly, you should be able to talk openly about training runs that
didn’t work
: what failed, how you diagnosed it, what it cost, and what you changed afterward.
The wider engineering environment spans:
- Robot learning and reinforcement learning
- Video and multimodal foundation models
- Perception for difficult real-world environments
- Custom physics simulation
- Rendering and light transport
- GPU kernel optimization
- Hardware-software co-design
That creates an unusually broad technical surface area. Your models won’t exist purely to improve benchmark scores. The longer-term objective is intelligence that can operate through physical systems and contribute to genuinely autonomous manufacturing.
The company is onsite in the Bay Area, with its R&D operation to be based in San Jose.
If you’ve personally taken large multimodal or video models through expensive pre‑training runs, including the painful ones, this is one of the more unusual opportunities to apply that experience to physical‑world AI.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).