AI Performance Engineer
Sunnyvale, Santa Clara County, California, 94087, USA
Listed on 2026-09-04
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Software Engineer
Applied Intuition, Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion, the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving machine on the planet. Applied Intuition services the automotive, defense, trucking, construction, mining and agriculture industries in three core areas: tools and infrastructure, operating systems, and autonomy.
Eighteen of the top 20 global automakers, as well as the United States military and its allies, trust the company’s solutions to deliver physical intelligence. Applied Intuition is headquartered in Sunnyvale, California, with offices in Washington, D.C.;
San Diego;
Ft. Walton Beach, Florida;
Ann Arbor, Michigan;
London;
Stuttgart;
Munich;
Stockholm;
Bangalore;
Seoul; and Tokyo. Learn more at
We are an in-office company, and our expectation is that full-time employees primarily work from their Applied Intuition office 5 days a week. However, we also recognize the importance of flexibility and trust our employees to manage their schedules responsibly. This may include occasional remote work, starting the day with morning meetings from home before heading to the office, or leaving earlier when needed to accommodate family commitments.
This in-office expectation does not apply to contractor positions
About the Role
We are looking for a performance engineer who specializes in making large-scale machine learning workloads fast and cost-efficient in the datacenter. This role is focused on distributed training runs spanning many nodes, and high-throughput batch inference sweeping petabytes of real-world autonomy logs for auto-labeling, data mining, ground-truth generation, and evaluation.
The optimization target here is not tail latency on a vehicle - it is throughput, cluster goodput, and cost per unit of data processed. A training run that wastes 30% of its GPU-hours on stalled data loaders, or an offline inference sweep that takes a week instead of a day, directly slows down how fast the whole company can iterate. You will own the gap between what our fleet of accelerators is theoretically capable of and what our workloads actually achieve: profiling across the stack, finding where the compute and the wall-clock time actually go, and closing the difference.
You will work at the intersection of accelerators, ML frameworks, and large-scale data infrastructure, partnering with the teams who own each layer to land wins that show up in training time-to-result and offline processing cost. At Applied, we encourage all engineers to take ownership over technical and product decisions, closely interact with users to collect feedback, and contribute to a thoughtful, dynamic team culture.
At Applied, you will:
Profile and optimize distributed training end to end - data loading and preprocessing, augmentation, kernel execution, gradient communication, and checkpointing
Optimize large-scale offline and batch inference over petabyte-scale sensor logs: batching and scheduling strategies, quantization and low-precision execution, graph optimization, and accelerator saturation across long-running sweeps
Establish roofline and performance models for our workloads, quantify the gap between achieved and theoretical performance, and stack-rank optimization opportunities by impact and effort
Improve multi-node scaling efficiency: sharding and parallelism strategies, collective communication, interconnect utilization, and memory-bandwidth and kernel-fusion bottlenecks
Drive cluster goodput - reduce GPU idle time from input pipeline stalls, storage and network I/O, scheduling gaps, stragglers, and failure recovery on long-running jobs
Build the benchmarking, observability, and regression-detection tooling that keeps performance from silently degrading as models and code evolve
Collaborate with engineers across functions to solve complex data and compute problems at scale
Contribute to a team culture that values effective collaboration, technical excellence, and innovation
Hands-on ML performance engineering experience: profiling, roofline analysis, throughput optimization, and root-cause investigation in production systems
Experience with distributed multi-node training at scale (FSDP, Deep Speed, Megatron, NCCL, or equivalent), including diagnosing scaling inefficiency as node count grows
Deep familiarity with GPU or accelerator performance concepts - memory bandwidth, kernel launch overhead, occupancy, quantization, collective communication
Experience with high-throughput or batch inference systems (NVIDIA Triton Inference Server, TensorRT, ONNX Runtime, Ray, or similar)
Fluency in Python and proficiency in C++ or another systems language
Excellent debugging, analytical, and problem-solving skills
A deep understanding of machine learning foundations, and the ability to develop technical solutions for problems with no established playbook
Nice to have:
GPU kernel development experience: CUDA,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).