Senior Software Engineer, ML Infrastructure Austin, TX
Listed on 2026-09-09
-
Software Development
Machine Learning/ ML Engineer, Software Engineer, AI Engineer (Applied/Software), DevOps
Senior Software Engineer, ML Infrastructure
Austin, TX
Apptronik is a human-centered robotics company developing AI-powered robots to support humanity in every facet of life. Our flagship humanoid robot, Apollo, is built to collaborate thoughtfully with people, starting with critical industries such as manufacturing and logistics, with future applications in healthcare, the home, and beyond.
We operate at the cutting edge of Applied AI, applying our expertise across the full robotics stack to solve some of society's most important problems. You will join a team dedicated to bringing Apollo to market at scale, tackling the complex challenges like safety, commercialization, and mass production to change the world for the better.
JOB SUMMARYApptronik is building Apollo, a general-purpose humanoid robot, and the physical AI that drives it. Scale is the name of the game: every robot and teleoperator we field produces synchronized video, proprioceptive, tactile, and force-torque streams, and the fleet's output grows with every deployment. Turning that volume of data into shipped autonomy — routinely, at multi-terabyte scale — is what this role is about.
We are looking for a Senior Software Engineer, ML Infrastructure to build that platform: the self-serve services and pipelines that carry data from collection through curation, training, and evaluation to a qualified model running on real hardware. Much of it is being created ground-up for the long term — humanoid robotics has few off-the-shelf answers — so the team builds first-party platform services alongside the open-source and commercial tooling we adopt where it genuinely fits.
This is a hands-on role on a small team whose platform is depended on daily by researchers and engineers across MLOps, Autonomy, Data Platform, and Tele Op.
ESSENTIAL DUTIES AND RESPONSIBILITIESYou will build the ML platform — the APIs, workers, and control planes that let researchers and robot teams move data and models through the system in a self-serve manner, with the testing and observability that being a dependency implies. The platform's responsibilities include:
- Data Curation & Annotation: Turn raw robot and simulation data into training-ready datasets — selection and filtering of manipulation episodes with synchronized sensor streams; annotation workflows that combine automatic labeling with human-in-the-loop review at throughput; and dataset versioning and lineage strong enough that any model traces back to the exact data that produced it.
- Data Pipelines at Scale: Make multi-terabyte dataset operations routine — transformation and assembly, coverage and quality statistics that tell us a training set is good before we spend a cluster-week on it, and read paths that keep GPUs fed.
- Simulation & Evaluation: Build the rollout harnesses that evaluate policies in simulation on our GPU cluster; the benchmarks and metrics captured consistently across simulation, real-robot, and teleoperation sources; and the qualification gates a model must pass before it reaches Apollo — automatic, not manual review.
- Model Promotion: Build the model store — versioning, metadata, attached evaluation results, lineage — and the promotion path from trained to qualified to deployed on robot, including packaging (ONNX, TensorRT) in partnership with Autonomy.
- Developer
Experience:
Provide the tooling researchers use daily — experiment tracking, training job submission, sweeps, and reproducible container environments. Reduce time from idea to running training job; win adoption by being the fastest path, not by mandate.
Alongside the technical work, you will partner with Autonomy, Data Platform, and Tele Op on dataset and model lifecycle contracts, contribute to the technical direction of these layers, and mentor the engineers around you through code and design review.
SKILLS AND REQUIREMENTSNo single person will have depth in everything below. We are looking for someone who has built platform services in production at scale with real depth in at least one of three areas —
large-scale data pipelines
, annotation and labeling
, or evaluation and simulation — plus solid cloud and Python across the board:
- A…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).