Senior Quantized Inference Engineer - Accelerate LLMs
Job in
Redmond, King County, Washington, 98052, USA
Listed on 2026-07-30
Listing for:
NVIDIA AI
Full Time
position Listed on 2026-07-30
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, Software Engineer
Job Description & How to Apply Below
NVIDIA is seeking a Senior Software Engineer for Quantized Inference to accelerate the development of efficient inference recipes for LLMs. You will implement quantized and sparse recipes, work on kernel and model-level implementations, and collaborate with partner teams to optimize throughput and interactivity across Megatron-LM, Model Opt, and vLLM.
The role requires strong Python skills with familiarity in C++, experience with PyTorch internals, and 4+ years in software engineering.
#J-18808-LjbffrPosition Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×