LLM Inference Engineer: Scale & Optimize Production
Listed on 2026-10-08
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Hippocratic AI is seeking an experienced LLM Inference Engineer to optimize its LLM serving infrastructure in Palo Alto, CA. You will design multi-node architectures, implement multi-LoRA serving, and apply quantization techniques to reduce footprint while maintaining quality.
The role emphasizes practical production deployment, benchmarking, and latency-focused optimizations across GPUs, with a strong emphasis on Python, C++, and CUDA.
We are currently recruiting a LLM Inference Engineer:
Scale & Optimize Production for our team in Palo Alto, CA, United States.
All applications are reviewed carefully by our team.
The position is based in Palo Alto, CA, United States.
This opportunity is part of our work in IT & Technology.
The advertised compensation is 150..
We aim to respond to suitable candidates as soon as possible.
Full responsibilities and requirements are described in the listing above.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).