Senior Solutions Architect, AI Performance Engineering
Listed on 2026-07-23
-
Software Development
AI Engineer (Applied/Software), Software Engineer, Unix/Linux
Overview
We are looking for a Solutions Architect with a performance engineering background who can help our most sophisticated Autonomous Vehicles and Robotics customers accelerate Physical AI workloads using NVIDIA's full-stack technologies.
As part of the Automotive Solutions Architecture team, we work with some of the most innovative accelerated computing platforms focused on the development and test of Autonomous Vehicles.
We dive deep into customer projects to solve performance bottlenecks and use insights from workloads to guide next-generation NVIDIA hardware and software.
Applications for this job will be accepted until July
212026.
We work in a diverse environment and are proud to be an equal‑opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
Responsibilities- Engage directly with key AV application engineers to understand current and future problems.
- Develop and improve fundamental parallel algorithms and data structures.
- Provide efficient GPU‑based solutions through library development and direct application contributions.
- Collaborate with the architecture, research, libraries, tools, and system software teams at NVIDIA to influence the design of next‑generation architectures, software platforms, and programming models.
- Perform deep performance optimization of high‑performance operators, including GPU kernel optimization, instruction‑level tuning, and compiler optimization.
- Support customers using NVIDIA libraries such as cuDNN, cuBLAS, and CUTLASS, as well as open‑source libraries like DeepGEMM, FlashMLA, Flash Attention, Flash infer, etc.
- Improve communication for Physical AI‑related distributed transformer workloads using NVIDIA communication tools (NCCL, NCCLGIN, NVSHMEM) and open‑source solutions (DeepEP, NCCLEP).
- Study interconnect topologies (NVLink) and network protocols (Infini Band/RoCE) to develop efficient data‑transfer strategies and enable compute‑communication overlap.
- BSc, MSc, PhD or equivalent experience in Computer Science, Electrical Engineering, Physics, Mathematics, or related field.
- 8+ years of validated ML/DL performance‑engineering experience improving GPU compute efficiency of training and inference workloads.
- Proficient in C, C++ or Python and Linux environments.
- Strong understanding of software development, programming techniques, and algorithms.
- Excellent mathematical fundamentals in linear algebra and numerical methods.
- Experience in parallel programming and high‑performance computing; knowledge of parallel architectures and performance‑analysis techniques.
- Experience in distributed communication optimization, including RDMA, GPU interconnects, collective communication algorithms, and associated open‑source libraries.
- Effective verbal and written communication and technical‑presentation skills; ability to communicate ideas/code through blog posts, Git Hub, PPT.
- Experience writing CUDA kernels, using Nsight System and Nsight Compute.
- Experience evaluating and improving full‑stack systems in LLM and HPC domains, from operator‑level through framework‑ to algorithm‑level optimizations.
- Proven software‑engineering fundamentals and system‑architecture thinking with capacity to build modules and lead engineering approaches in complex systems.
Base salary ranges are USD
184,000–287,500 for Level4 and USD
224,000–356,500 for Level
5. Additional equity and benefits will also be provided.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).