×
Register Here to Apply for Jobs or Post Jobs. X

Senior AI Infrastructure Engineer

Job in Costa Mesa, Orange County, California, 92626, USA
Listing for: Anduril
Full Time position
Listed on 2026-10-05
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, DevOps
Salary/Wage Range or Industry Benchmark: 150000 - 230000 USD Yearly USD 150000.00 230000.00 YEAR
Job Description & How to Apply Below

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century's most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center.

As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.

ABOUT THE TEAM

The Air Dominance & Strike team at Anduril develops aerial and multi-domain robotic systems. The team is responsible for taking products like Fury (unmanned fighter jet) and Barracuda (air-breathing cruise missile) from concept to product. The team also develops Lattice for Mission Autonomy, Anduril's premier software platform that enables masses of Fury, Barracuda, and other first and third party robots to collaborate across various missions.

We work in close coordination with specialist teams like Perception, Motion Planning, Hardware, and Test Engineering to solve some of the hardest problems facing our customers. We are looking for software engineers and roboticists excited about creating a powerful autonomy software stack that includes computer vision, motion planning, SLAM, controls, estimation, and secure communications.

ABOUT THE JOB

We are looking for a Senior AI Infrastructure Engineer to build, scale, and optimize the end-to-end machine learning platform that powers Anduril's autonomous systems.

In this role, you will own critical components of our ML platform and MLOps tooling. You will build and operate the infrastructure required to train, evaluate, host, and serve complex AI models (including LLMs, computer vision, and RL agents) across cloud environments and air-gapped, edge-deployed networks. Working closely with AI Research Scientists and Platform Engineers, you will eliminate friction in model development, optimize hardware utilization, and ensure robust delivery of models into safety-critical operational environments.

WHAT

YOU'LL DO
  • Build, optimize, and maintain scalable training, orchestration, and experimentation infrastructure to accelerate state-of-the-art model development.
  • Identify and resolve bottlenecks in the ML lifecycle by developing tooling for experiment tracking, automated profiling, and hyperparameter tuning.
  • Implement and scale robust data pipelines (ETL) to process multi-modal data (video feeds, radar, flight telemetry, and simulation logs) captured from physical assets and test sites.
  • Deploy high-throughput, low-latency model serving frameworks optimized for both cloud environments and resource-constrained, air-gapped tactical edge hardware.
  • Develop robust CI/CD pipelines for ML models, including automated regression testing, validation benchmarks, and safe rollout/rollback strategies.
  • Implement pipelines for model evaluation, validation, and reinforcement learning alignment loops (RLHF/DPO) to ensure predictability and safety in mission-critical deployments.
  • Partner with AI Researchers and Computer Vision engineers to translate modeling requirements into scalable, reusable infrastructure.
  • Mentor peers, conduct thorough design and code reviews, and champion engineering best practices across the team.
REQUIRED QUALIFICATIONS
  • 5+ years of software engineering experience with demonstrated success in building and operating production-scale machine learning infrastructure or distributed systems.
  • Proficiency in Python, Go, or C++, with a strong grasp of software engineering fundamentals, systems design, and concurrent programming.
  • Hands-on experience with container orchestration (Docker, Kubernetes) and distributed training frameworks (e.g., PyTorch Distributed, Ray, Slurm, or Megatron-LM).
  • Experience building and maintaining distributed data pipelines handling large-scale unstructured or multi-modal datasets.
  • Track record of owning projects end-to-end-from technical design to production deployment and operational monitoring.
  • Eligible to obtain and maintain an active U.S. Top Secret security clearance.
PREFERRED QUALIFICATIONS
  • Experience deploying ML infrastructure, model serving, or artifacts in secure, air-gapped, or regulated environments (e.g., IL5/IL6, Gov Cloud).
  • Hands-on…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary