×
Register Here to Apply for Jobs or Post Jobs. X

Machine Learning Infrastructure Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Atoms
Full Time position
Listed on 2026-07-08
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, Data Engineering, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 224000 - 280000 USD Yearly USD 224000.00 280000.00 YEAR
Job Description & How to Apply Below
Position: Staff Machine Learning Infrastructure Engineer

Who we are

Atoms is building the machines that power the next era of progress.

Over the last decade, software has transformed the digital world. But the physical world, where food is made, minerals are mined, goods are moved, and industries are run, remains far less intelligent, far less efficient, and far more constrained. We’re changing that.

Atoms builds Physical AI – real‑world robots for the industries that move civilization forward, starting with food, mining, and transport. Our systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive.

This work requires more than robotics. It requires deep integration across hardware, software, AI, operations, manufacturing, and real estate. We don’t just build machines in a lab. We deploy them into real environments, operate them, learn from them, and improve them until they work at scale.

We are roboticists, engineers, operators, and builders. We believe the next great technology companies will not only transform information, but the physical systems that shape everyday life.

If you want to work on hard problems with real‑world impact, join us.

What you’ll do

We are seeking a foundational Machine Learning Infrastructure Engineer to design and build the large‑scale ML training infrastructure that powers our next‑generation autonomous transport models. In this role, you will design the high‑performance training pipelines and validation environments that enable our world‑class robotics and ML researchers to iterate rapidly. You will own the challenge of scaling distributed GPU workloads to support a high volume of concurrent training runs across an expanding vehicle fleet, building a platform that can flexibly run on whatever GPU capacity is available, regardless of provider or environment, directly accelerating innovation across the platform.

  • Training Infrastructure:
    Design, implement, and scale repeatable machine learning infrastructure utilizing Kubernetes to support large‑scale distributed GPU training of novel neural networks.
  • Distributed Computing & Orchestration:
    Leverage distributed compute frameworks to efficiently manage and execute a high volume of complex ML training jobs concurrently across large GPU clusters.
  • Experiment Tracking & MLOps:
    Integrate advanced model management and experiment tracking tools to provide researchers with deep observability into training metrics and run performance.
  • Data Engineering Pipelines:
    Build and optimize high‑throughput data ingestion pipelines to seamlessly stream petabyte‑scale multi‑sensor vehicle logs into training environments.
  • Validation at Scale:
    Architect robust infrastructure for autonomous model validation and continuous integration testing, ensuring new vehicle policy releases are entirely regression‑free.
  • Cross‑Functional

    Collaboration:

    Partner closely with core robotics engineers and machine learning researchers to eliminate workflow bottlenecks and accelerate the deploy‑to‑vehicle lifecycle.
What we’re looking for
  • 8+ years of professional software engineering career experience
  • Strong backend systems programming skills with proficiency in Go, Python, Java or similar (with familiarity or exposure to Rust considered a plus).
  • Proficiency with Kubernetes for container orchestration and building cloud‑agnostic environments from scratch.
  • Experience implementing distributed ML compute frameworks (e.g., Ray) to coordinate large pools of GPUs for heavy, multi‑node workloads.
  • Hands‑on experience building MLOps pipelines, metadata tracking architectures, and model registries using platforms like MLflow.
  • Prior experience managing high‑throughput data pipelines using modern distributed data engines to feed data‑hungry neural network architectures.
Why join us

At Atoms, you’ll work on one of the defining challenges of our time – bringing automation into the physical world to drive real, lasting impact. We exist to uncover valuable unknown truths and turn them into progress, which means constantly pushing beyond what’s known and building what doesn’t yet exist. The work is ambitious and often…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary