×
Register Here to Apply for Jobs or Post Jobs. X

Senior Solutions Architect, AI Performance Engineering

Job in Santa Clara, Santa Clara County, California, 95053, USA
Listing for: NVIDIA AI
Full Time position
Listed on 2026-07-24
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Software Engineer
Salary/Wage Range or Industry Benchmark: 184000 USD Yearly USD 184000.00 YEAR
Job Description & How to Apply Below

Job Requisition

Job Category:
Sales

Time Type:
Full time

We are looking for a Solutions Architect with a performance engineering background who can help our most sophisticated Autonomous Vehicles and Robotics customers accelerate Physical AI workloads using NVIDIA's full‑stack technologies. As part of the Automotive Solutions Architecture team, we work with some of the most innovative accelerated computing platforms focused on the development and testing of Autonomous Vehicles. We dive deep into customer projects to solve performance bottlenecks and use insights from workloads to guide next‑generation NVIDIA hardware and software.

If you are driven by innovation and ambition, this is the team for you.

What You’ll Be Doing
  • Engage directly with key AV application engineers to understand current and future problems and develop and improve fundamental parallel algorithms and data structures. Provide efficient solutions using GPUs through library development and direct application contributions.
  • Collaborate closely with the architecture, research, libraries, tools, and system software teams at NVIDIA to influence the build of next‑generation architectures, software platforms, and programming models.
  • Engage in deep optimization of high‑performance operators, including GPU kernel optimization, instruction‑level tuning, and compiler optimization. These optimizations directly support customers using NVIDIA libraries such as cuDNN, cuBLAS, and CUTLASS and open‑source libraries like DeepGEMM, FlashMLA, Flash Attention, Flash infer, etc.
  • Improve communication for Physical AI‑related distributed transformer workloads by employing NVIDIA communication tools such as NCCL, NCCL GIN, NVSHMEM, and open‑source solutions like DeepEP and NCCL EP. Conduct in‑depth studies of interconnect topologies (NVLINK) and network protocols (Infini Band/RoCE) to develop efficient data transfer strategies and enable compute‑communication overlap.
What We Need To See
  • BSc/MSc/PhD or equivalent experience in Computer Science, Electrical Engineering, Physics, Mathematics, or a related technical field.
  • 8+ years of hands‑on validated ML/DL performance engineering experience focused on improving GPU compute efficiency of training and inference workloads.
  • Experience with C, C++, or Python and proficiency with Linux.
  • Solid understanding of software development, programming techniques, and algorithms.
  • Strong mathematical fundamentals, including linear algebra and numerical methods.
  • Background in parallel programming and high‑performance computing, with extensive knowledge of parallel architectures and performance analysis and tuning. Experience in GPU programming is desirable.
  • Experience in distributed communication optimization is highly helpful, including familiarity with remote direct memory access, GPU interconnects, collective communication algorithms, and associated open‑source libraries used in large‑scale model training and inference.
  • Effective verbal and written communication and technical presentation skills. Ability to communicate ideas and code clearly through blog posts, Git Hub, and presentations.
Ways To Stand Out From The Crowd
  • Prior experience in writing CUDA kernels and using Nsight System and Nsight Compute.
  • Experience evaluating and improving the full‑stack system in at least one of these areas: LLM and HPC. Expertise ranging from operator‑level through framework‑level to algorithm‑level optimization is strongly preferred.
  • Proven software engineering fundamentals and system architecture thinking, with the ability to build modules and lead engineering approaches in complex systems.

Salary for Level 4: 184,000 USD – 287,500 USD;
Level 5: 224,000 USD – 356,500 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted until July 21, 2026.

NVIDIA is committed to fostering an inclusive work environment and is proud to be an equal‑opportunity employer. NVIDIA does not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status, or any other characteristic protected by law.

#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary