×
Register Here to Apply for Jobs or Post Jobs. X

ML Performance Engineer; Training Efficiency

Job in Sunnyvale, Santa Clara County, California, 94086, USA
Listing for: Wayve
Full Time, Apprenticeship/Internship position
Listed on 2026-07-01
Job specializations:
  • IT/Tech
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Systems Engineer
Salary/Wage Range or Industry Benchmark: 336400 - 359000 USD Yearly USD 336400.00 359000.00 YEAR
Job Description & How to Apply Below
Position: Staff ML Performance Engineer (Training Efficiency)

The Role

We are looking for a Staff ML Performance Engineer to join our Training Tech team working on optimizing large scale ML jobs to enable scaling our models to the next order of magnitude. A successful candidate will increase efficiency of training and inference workloads in order to allow Wayve to train larger models faster.

Key responsibilities:

  • Profile ML workloads to identify their bottlenecks, e.g. using NVIDIA Nsight Systems
  • Design and implement efficiency improvements to maximize MFU and throughput, e.g. parallelism, model compilation, mixed precision
  • Design and implement observability tools to identify bottlenecks and drive performance improvements, e.g. to track MFU, throughput, latency, etc
  • Design and implement benchmarking tools, e.g. to track efficiency gains or regressions
  • Collaborate closely with Research teams to integrate training efficiency improvements and create a culture of performance optimization
About You

In order to set you up for success in this role, we're looking for the following skills and experience.

Essential

  • 10+ years of industry experience driving performance engineering across ML systems, GPU compute infrastructure, distributed platforms or similar field.
  • Experience optimizing large scale jobs on GPU compute clusters.
  • Experience in working in platform teams and working with research teams.
  • Experience in writing, reporting, and tracking performance benchmarks in an open and accessible way.
  • Ability to write high quality, well-structured and tested Python code
  • BS or MS in Machine Learning, Computer Science, Engineering, or a related technical discipline or equivalent experience

Desirable

  • Experience working with concurrent, parallel and distributed computing.
  • Experience using NVIDIA NSight Systems or other system profilers.
  • Experience implementing GPU kernels (CUDA, Triton, etc).
  • Knowledge of computing fundamentals - what makes code fast, secure and reliable.

This role is a full-time role based in Sunnyvale, CA (hybrid) and the reasonably estimated salary for this role ranges from $336,400 to $359,000, plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience.

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary