Sr. Principal Software Engineer
Job in
Burlington, Middlesex County, Massachusetts, 01805, USA
Listed on 2026-07-04
Listing for:
Cerence Inc.
Full Time
position Listed on 2026-07-04
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Software Engineer, Machine Learning/ ML Engineer, DevOps
Job Description & How to Apply Below
We use cookies to understand how you use our site and to improve your experience. This includes personalizing content and advertising. To learn more, . By continuing to use our site, you accept our use of cookies and Privacy Policy#Sr. Principal Software Engineer page is loaded## Sr. Principal Software Engineer Apply locations:
Remote - USAtime type:
Full time posted on:
Posted Yesterday job requisition :
R0005965##
** A Moving Experience.
**** Who is Cerence AI?
** Cerence AI is the global leader in AI for transportation, specialized in building AI and voice-powered companions for cars, two-wheelers, and more that enable people to focus on what matters most. With over 500 million cars shipped with Cerence AI's technology, we partner with leading automakers (such as Volkswagen, Mercedes, Audi, Toyota and many more), mobility providers, and technology companies to power intuitive, integrated experiences that create safer, more connected, and more enjoyable journeys for drivers and passengers alike.
** Our Driving Force
** Our team is dedicated to pushing the boundaries of AI innovation, working around the globe with headquarters in Burlington, Massachusetts, USA and 16 other offices across Europe, Asia, and North America. We bring together diverse backgrounds, and varied skill sets with the shared goal of advancing the next generation of transportation user experiences. Our culture is customer-centric, collaborative, fast-paced, and fun, with continuous opportunities for learning and development to support your career growth.
Interested in having a significant impact in a dynamic industry with a high-performing global team? We’re looking for an exceptional Senior Principal Software Engineer who is ready to drive the future of mobility with us!
*
* Job Description:
**** What You Will Work On
*** Optimize and deploy
** high**‐
** performance LLM inference pipelines
*** Own inference runtimes across
** data center, edge, and embedded platforms
*** Push model performance through quantization, kernel fusion, and cache optimization
* Drive latency and throughput improvements that directly impact production products
* Enable efficient, reliable deployment without external vendor dependency
** Core Responsibilities
*** Inference Engines & Runtime
** Build deep expertise and ownership of:
* vLLM* TensorRT‐LLM
* llama.cpp
* QAIRT* Extend and tune inference engines using
** custom CUDA kernels
*** Adapt runtimes for constrained and embedded deployment environments
*** Quantization & Numerical Optimisation
**** Implement and evaluate quantisation strategies:
* INT8, INT4, FP4, FP8, mixed precision
* AWQ* GPTQ
* Balance accuracy, latency, memory footprint, and throughput
*** KV Cache Optimization
**** Optimize key–value cache performance through:
* Paging* Prefix caching
* Cache‐aware memory layout design
* Reduce memory pressure while sustaining high throughput
*** Latency & Throughput Optimisation
**** Design and tune:
* Batching strategies
* Continuous batching
* Speculative decoding
* Optimize tail latency and tokens/sec under real production traffic patterns
** What Success Looks Like
*** Models deploy efficiently on edge and embedded devices, not just servers
* Tokens/sec significantly outperform baseline implementations
* End‐to‐end latency is minimized and predictable
* Inference cost per request is materially reduced
* The company is no longer dependent on partners for inference optimization
** Required Experience & Skills
***** Strongly Required
**** Proven experience optimizing
** ML inference performance in production
*** Deep understanding of
** GPU architecture and memory hierarchies
*** Hands‐on experience with CUDA and low‐level performance tuning
* Experience deploying models beyond research environments
*** Critical Technical Skills
**** Inference engines: vLLM, TensorRT‐LLM, llama.cpp, QAIRT
* CUDA kernel development and profiling
* Quantisation techniques: INT8/INT4/FP4/FP8, AWQ, GPTQ
* KV cache optimisation and memory layout design
* Latency optimisation: batching, speculative decoding, continuous batching
** Common Problems You’ll Be Solving
*** Deploy efficiently on edge or embedded targets
*…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×