×
Register Here to Apply for Jobs or Post Jobs. X

Principal Software Engineer – E2E Performance, Goodput

Job in California, Moniteau County, Missouri, 65018, USA
Listing for: Jobtailor
Full Time position
Listed on 2026-08-02
Job specializations:
  • Software Development
    AI Reliability/ Performance Engineer
Salary/Wage Range or Industry Benchmark: 150000 - 210000 USD Yearly USD 150000.00 210000.00 YEAR
Job Description & How to Apply Below
Location: California

  • Drive performance characterization work streams with engineering teams of key CSP/hyperscale customers — ensuring they understand platform performance expectations, profiling methodology, and tuning options for their specific workloads
  • Gather and synthesize CSP performance feedback — identify gaps between expected and actual throughput, and champion optimization priorities back into NVIDIA's CUDA, NCCL, driver, and firmware teams
  • Ensure key open-source performance and stress tools (e.g., STREAM, GPU Burn, GPU BLAST) are updated and validated for the latest NVIDIA rack‑scale systems, GPU architectures, and CPU platforms — so customers and internal teams have reliable baseline measurements from day one
  • Work closely with CSPs to ensure their own performance and validation tooling reflects the latest GPU capabilities, memory hierarchy changes, and platform‑specific tuning parameters
  • Conduct cross‑CSP performance comparison and pattern analysis — identify configuration, software, or workload differences that explain performance gaps between deployments
  • Collaborate with CSPs to ensure performance‑related integration work (profiling infrastructure, benchmark harnesses, config validation) is ready ahead of deployment milestones
  • Define test strategies and tooling requirements for performance validation — both for NVIDIA internal certification and customer acceptance
Requirements
  • 15+ years of experience in systems performance engineering, ideally in GPU/HPC/ML infrastructure.
  • BS or MS in Computer Science, Computer Engineering, or related field (or equivalent experience)
  • Proficiency in GPU workload profiling: nsight systems, nsight compute, DCGM metrics, or equivalent instrumentation
  • Understanding of distributed training performance dynamics: computation/communication overlap, pipeline bubbles, memory bandwidth utilization, collective efficiency
  • Statistical methods for performance analysis: regression detection, confidence intervals, A/B comparison at scale
  • Understanding of how the full software stack impacts performance: driver overhead, collective algorithm selection, memory allocation, scheduling, firmware power management
  • Strong data analysis and visualization skills (Python, pandas, dashboards).
  • Customer obsession — genuine passion for understanding why customers aren't achieving expected performance and driving solutions
  • Ability to communicate performance findings to both deep technical audiences and executive leadership
  • Demonstrated success influencing multiple engineering teams to prioritize performance improvements
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary