×
Register Here to Apply for Jobs or Post Jobs. X

Technical Lead - Autonomy Evaluation

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Atoms
Full Time position
Listed on 2026-10-08
Job specializations:
  • IT/Tech
Salary/Wage Range or Industry Benchmark: 185000 - 242000 USD Yearly USD 185000.00 242000.00 YEAR
Job Description & How to Apply Below

Atoms is building the machines that power the next era of progress.

Over the last decade, software has transformed the digital world. But the physical world, where food is made, minerals are mined, goods are moved, and industries are run, remains far less intelligent, far less efficient, and far more constrained. We’re changing that.

Atoms builds Physical AI— real-world robots for the industries that move civilization forward, starting with food, mining, and transport. Our systems are designed to understand, predict, and control the real world with precision, turning complex physical operations into something more reliable, more scalable, and more productive.

This work requires more than robotics. It requires deep integration across hardware, software, AI, operations, manufacturing, and real estate. We don’t just build machines in a lab. We deploy them into real environments, operate them, learn from them, and improve them until they work at scale.

We are roboticists, engineers, operators, and builders. We believe the next great technology companies will not only transform information, but the physical systems that shape everyday life.

If you want to work on hard problems with real-world impact, join us.

About the role

We are seeking a Technical Lead for Autonomy Evaluation to own the metrics our autonomy releases are judged on and the log replay and evaluation platform that computes them. In this role, you will take evaluation from recorded logs, to scored regression suites running on every candidate release, to the report a release passes before it reaches a vehicle, developing how replay, scoring, metrics, and test coverage come together into one platform engineers use every day.

You will do this across on-road vehicles, and you will build the evaluation infrastructure that lets each new software release, sensor configuration, and platform be assessed seamlessly.

What you’ll do
  • Metrics. Own the safety and behavior metrics a release is judged on, including collision and near-miss measures, trajectory agreement against ground truth, and comfort, and the go/no-go release criteria built on them.
  • Log replay and scoring. Own the pipeline that converts recorded logs into scored test cases, covering open-loop and closed-loop evaluation of perception, localization, and planning against ground truth.
  • Regression and benchmarking. Own the system that benchmarks new software against old across the log corpus for every candidate release, including evaluation dataset design, regression detection, and triage of a regressed case to an owner.
  • Test coverage and data mining. Own how the corpus is mined for the long tail of rare and important events, including sampling strategy, selection of logs by expected value per replay-hour, and coverage of scenario classes the corpus does not yet contain.
  • Execution  deterministic and reproducible execution of evaluation, run identity, traceability of every result to a software version, and cost per replay-hour as a tracked number.
  • Engineering practices. Set the practices for metric definitions, test design, code review, and the write-up of evaluation results.
  • Technical standard. Mentor the engineers who join the function, set the technical bar for their work, and participate in hiring.
What we’re looking for
  • Bachelor's degree in Computer Science, Robotics, Electrical Engineering, Statistics, or a related field with 8+ years of relevant experience.
  • 2+ years as a technical lead or Engineering manager for an evaluation, validation, or test team.
  • Ownership of an offline or system-level evaluation system for an autonomous vehicle or robotics program, or a comparable large-scale ML model evaluation system in production, with accountability for the metrics and for the release…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary