Senior/Staff Software Engineer, AI Agent Infrastructure
Job in
Mountain View, Santa Clara County, California, 94039, USA
Listed on 2026-08-04
Listing for:
NextGenEnergyJobs
Full Time
position Listed on 2026-08-04
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Software Engineer
Job Description & How to Apply Below
Our mandate is to amplify the output of every engineer and researcher at Nuro by 100x.
Key Responsibilities- We already operate a substantial agent system in production — a fleet of agents with an extensive library of skills and plugins, integrated into the tools our engineers use daily. This role is about what it takes to make that system trustworthy, autonomous, and an order of magnitude more capable. Three things sit at the center of it.
- Closed-loop evaluation. Our ambition is to build the most rigorous evaluation system for AI work anywhere — closed-loop, meaning every agent action produces a measurable outcome that feeds back into whether that agent is trusted to act again. Acceptance, revert, and override rates per workflow. Statistical honesty about whether a difference is real. Regression detection that fires before a human notices.
Everything else on this team depends on this being right, and almost nobody has built it well. - Agent platform. The runtime that makes autonomous agents safe to run against real systems: orchestration, sandboxing and isolation, tool and skill frameworks, memory, identity and permissioning, and the gateway and observability layer underneath. Agents that touch production code and production infrastructure need containment and auditability before they need capability.
- Auto research infrastructure. The automation of the research loop itself: agents that read the current state of a model and its metrics, form a hypothesis, launch an experiment, evaluate the result honestly, and either propose a change or discard the idea and move on. At Nuro that loop runs against the training pipelines behind the driving model — real experiments, real compute budgets, real metrics that determine whether a behavior ships.
The hard parts are trusting the measurement, surviving experiments that take days, spending finite research compute wisely, and producing proposals a skeptical researcher can audit and reject. - Alongside this, the team builds agent-powered tooling across the engineering lifecycle — code generation, review, debugging, test and CI failure attribution, knowledge retrieval, triage. There is also appetite on this team for post-training our own models where an internal workload justifies it, and the engineer in this role would be central to that work.
- About the Work What You Might Own in Your First Two Quarters
- Build the closed-loop measurement layer that tells us, per workflow, whether agent output is accepted, reverted, or overridden — and use it to decide where autonomy expands and where it gets pulled back.
- Take the autoresearch loop from assisted to unattended for a bounded class of experiments, including the eval and confidence machinery required to run it without a human in the loop.
- Design the isolation and permissioning model that lets agents act on production repositories and infrastructure with an auditable record of what they did and why.
- 5+ years of software engineering experience (or 4+ with a Master's) in computer science, engineering, or equivalent practical experience. Staff-level candidates should bring correspondingly deeper scope and ownership.
- Deep, current taste in LLM research. You understand how a model is trained from scratch — data, tokenization, architecture, pretraining dynamics, the full post-training stack of supervised fine-tuning, preference optimization, and RL — and you can reason about what a training decision does to model behavior. You follow the literature because you want to, not because it is on a roadmap.
- You know what happens under the hood ention and KV-cache behavior, batching and scheduling, quantization, speculative decoding, prefix caching, context handling, and how each trades off latency, throughput, and cost.
- You have built and operated LLM-based agent systems in production — tool use, orchestration, sandboxing, retrieval, memory — and you know where they break.
- Strong backend and distributed systems background at scale: cloud infrastructure, service design, storage, queuing, and the judgment to build things that stay up.
- Strong programming skills in Python.
- You are…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×