LLM Inference Engineer; Mid, Sr
Listed on 2026-10-04
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
As HAI's LLM Inference Engineer, you will own the serving infrastructure that determines whether our breakthrough healthcare AI reaches patients efficiently and reliably. You'll optimize the systems that translate raw model capability into sub-100ms responses-makings difference between conversational experiences that feel natural and those that feel broken. This role exists because inference optimization at scale is where research meets reality: your work directly determines latency, cost, and availability for millions of patient conversations across healthcare systems.
What You Will AccomplishOwn your first major outcome: By day 90, you will have shipped a measurable improvement to our inference serving stack (reduce latency, improve throughput, or optimize cost per inference), validated the gains across our production deployment scenarios, and established the performance optimization roadmap that will guide infrastructure investment.
Drive lasting impact: At 12 months, you will have designed and deployed advanced serving architectures (disaggregated inference, optimized caching, speculative decoding) that meaningfully improve patient experience and operational efficiency, contributed novel optimization techniques that become part of our core infrastructure, and made our serving stack a durable competitive advantage in healthcare AI deployment.
The TeamYou’ll work alongside systems engineers, ML researchers, and infrastructure experts who are obsessed with making AI systems fast, reliable, and cost-effective. This is a team that values deep technical rigor, continuous benchmarking, and solving hard systems problems that have real impact on patient experience and business unit economics.
What You’ll DoDesign and implement multi-node serving architectures for distributed LLM inference
Optimize multi-LoRA serving systems
Apply advanced quantization techniques (FP4/FP6) to reduce model footprint while preserving quality
Implement speculative decoding and other latency optimization strategies
Develop disaggregated serving solutions with optimized caching strategies for prefill and decoding phases
Continuously benchmark and improve system performance across various deployment scenarios and GPU types
We believe the best ideas happen together. This role is based in our Menlo Park, California office, expected to be in office five days a week.
CompensationCompensation is based on experience, expertise, and level of responsibility. We offer competitive packages that reflect the seniority and scope of the role, along with equity, health insurance, and other benefits.
What You BringMust-Have:
Experience optimizing LLM inference systems at scale
Proven expertise with distributed serving architectures for large language models
Hands-on experience implementing quantization techniques for transformer models
Strong understanding of modern inference optimization methods, including:
Speculative decoding techniques with draft models
Eagle speculative decoding approaches
Proficiency in Python and C++
Experience with CUDA programming and GPU optimization
Contributions to open-source inference frameworks such as vLLM, SGLang, or TensorRT-LLM
Experience with custom CUDA kernels
Track record of deploying inference systems in production environments
Deep understanding of performance optimization systems
Show us what you've built: Tell us about an LLM inference or training project that makes you proud! Whether you've optimized inference pipelines to achieve breakthrough performance, designed innovative training techniques, or built systems that scale to billions of parameters – we want to hear your story.
Open source contributor? Even better! If you've contributed to projects like vllm,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).