Senior Software Engineer, Quantized Inference
Listed on 2026-08-26
-
Software Development
Machine Learning/ ML Engineer, AI Engineer (Applied/Software), Python, Software Engineer
Implement quantized and sparse recipes within inference engines and manage model export pipelines to ensure correct serialization. Develop benchmarking harnesses, data analysis tools, and improve developer productivity through infrastructure and CI improvements.
Requirements:Requires proficiency in Python and familiarity with C++, along with strong software engineering fundamentals and experience with ML accelerators. Candidates should have experience with PyTorch internals and a minimum of 4 years in a relevant software engineering role, preferably with a MS/PhD in Computer Science.
KeySkills:
Python, C++, Triton Kernels, PyTorch, Quantized Inference, Model Compression, Machine Learning Accelerators, vLLM, TRT-LLM, SGLang, Megatron-LM, Model Opt, Software Engineering, Data Analysis, Numerical Debugging, Large Language Models
Benefits:- Equity
- Benefits
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).