AI Automation Engineer, World Test Lab
Listed on 2026-10-02
-
IT/Tech
IT QA Tester / Automation, AI Engineer (Applied/Software)
At Niantic Spatial, we're building the future of physical AI. Powered by a proprietary database of over 30 billion posed images, our groundbreaking mapping technology unlocks a new dimension of interaction and spatial intelligence that helps both humans and machines better understand, represent, navigate, and engage with the real environment.
Our reconstruction technology captures environments with geometric accuracy and extreme detail from any standard camera, and our Visual Positioning System delivers precise positioning almost anywhere in the world. We serve customers across robotics, the public sector, and energy and industrial markets - building for the 80% of economic activity that takes place beyond our screens.
About the Real-World Test LabPhysical AI doesn't get graded on a leader board. It gets graded on a factory floor at shift change, in a substation with no GPS, on a site where the lighting is wrong and the stakes are real.
The Real-World Test Lab closes the gap between the benchmark and the field. We bring the customer's world inside our walls - their devices, their environments, their hardest conditions, and their definition of success - and make it the bar every release has to clear.
As Niantic Spatial's first and most demanding customer, we push our reconstruction, localization, and spatial understanding to their limits, find where they shine and where they break, and turn that into evidence that shapes what we build next. It's a new team at the frontier of physical AI, and you'll help invent how the job is done.
About the RoleWe're hiring an AI Automation Engineer, reporting to the Director of the Real-World Test Lab, to build the system that produces our evidence. Today our evaluations are manual, inconsistent, and slow. You'll design and own the automation that runs them end to end on every relevant release, with no human driving it, turning one-off experiments into an always-on service the whole company relies on.
This role is about owning the evidence, not executing a test plan. You'll decide how each workflow gets exercised, and you'll be measured on whether the company can trust the results, not on how much automation exists. The work counts when it keeps running correctly months later, without you in the loop.
You believe evaluation is engineering, not process. You've built systems that test other systems, and you know the difference between a script that works on your laptop and infrastructure a team can trust. When a result couldn't be reproduced, you fixed the tooling instead of arguing about the number.
What You'll DoBuild the Evaluation Machine - Own the automation that executes Lab evaluations end to end: environment setup, run orchestration, artifact capture, and result collection. Make reruns free so we test constantly rather than occasionally.
Make Results Comparable - Instrument scorecards so a result can be compared across product versions, devices, and capture conditions. A number without its lineage is not evidence.
Automate the Agent-Driven Layer - Build the agent workflows that exercise priority customer use cases at realistic scale, across the graded difficulty spectrum from easy to frontier, and be honest about where agents cannot yet replace human judgment.
Kill Manual Work Permanently - Convert one-off experiments into standing protocols that run on every relevant release. Automate recurring inspection wherever it can be automated, and route what genuinely cannot to scalable human review.
Make Failures Actionable - Produce diagnostics precise enough that a finding reaches its owner with a reproducible case and data attached. Findings that need re-investigation before anyone can act on them are half-finished.
Keep Data From Being the Bottleneck - Work with the AI Data Manager so every run is reproducible from a known dataset state, without anyone downloading and re-uploading data by hand.
30 days: One priority workflow evaluated end to end with no manual steps, with results in a comparable scorecard.
60 days
:
Evaluations trigger automatically on relevant releases, with diagnostics that route failures to the right owners.90 days: Three priority workflows under standing automated evaluation, with version-over-version comparison for leadership.
- Built and maintained production-grade automation or test infrastructure that other engineers relied on daily.
- Experience evaluating systems where correctness is graded rather than binary,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).