More jobs:
Deep Learning Software Engineer, TensorRT - College Grad
Job in
Santa Clara, Santa Clara County, California, 95053, USA
Listed on 2026-07-23
Listing for:
NVIDIA Corporation
Full Time
position Listed on 2026-07-23
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI Reliability/ Performance Engineer
Job Description & How to Apply Below
Job Overview
NVIDIA is seeking an experienced Deep Learning Software Engineer, TensorRT Performance to analyze and improve the performance of NVIDIA’s inference ecosystem, including TensorRT, TensorRT‑EdgeLLM and Torch‑TensorRT.
Responsibilities- Establish groundbreaking performance benchmarking methodologies and analysis workflows and identify performance issues and opportunities for NVIDIA’s inference ecosystem (e.g. TensorRT, TensorRT‑EdgeLLM, Torch‑TensorRT).
- Contribute features and code to NVIDIA/OSS inference frameworks including but not limited to TensorRT, TensorRT‑EdgeLLM, Torch‑TensorRT.
- Develop new model pipelines for NVIDIA’s inference ecosystem with optimized performance including quantization, scheduling, memory management, and distributed inference to set the gold standard for Gen AI performance.
- Work with cross‑collaborative teams inside and outside of NVIDIA across generative AI, automotive, robotics, image understanding, and speech understanding to set directions and develop innovative inference solutions.
- Scale performance of deep learning models across different architectures and types of NVIDIA accelerators.
- Bachelor’s, Master’s, PhD, or equivalent experience in Computer Science, Computer Engineering, EECS, or AI.
- At least 2 years of relevant software development experience.
- Strong C++ and Python programming and software engineering skills.
- Experience with deep learning frameworks (PyTorch, JAX, Tensor Flow, ONNX) and inference libraries (TensorRT, TensorRT‑LLM, vLLM, SGLang, Flash Infer).
- Experience with performance analysis and performance optimization.
- Strong foundation and architectural knowledge of GPUs.
- Deep understanding of modern deep learning models and workloads (Transformers, Recommenders, ASR, TTS, Visual Understanding).
- Proficiency in at least one deep learning programming domain‑specific language (CUDA, TileIR, CuTeDSL, cutlass, Triton).
- Prior contributions to major LLM inference frameworks (e.g. vLLM) or experience with graph compilers in deep learning inference (e.g. Torch Dynamo, Torch Inductor).
- Prior experience optimizing performance for low‑latency, resource‑constrained systems or embedded AI pipelines (e.g. Jetson systems, other edge AI accelerators).
Base salary ranges:
Level2: $124,000–$195,500 USD;
Level3: $152,000–$241,500 USD. Eligible for equity and benefits.
Applications will be accepted until at least July
24,2026.
NVIDIA is committed to fostering an inclusive work environment and is an equal opportunity employer. We do not discriminate on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×