Remote Inference Runtime Engineer LLM & Multimodal AI
Washington, District of Columbia, 20022, USA
Listed on 2026-09-30
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI Reliability/ Performance Engineer
Inferact is seeking an inference runtime engineer to advance LLM and diffusion model serving. You will optimize how models execute across diverse hardware, shaping the core of vLLM and enabling faster AI inference.
This remote role embraces flexible timezones with Pacific overlap for critical syncs, and compensation includes salary plus equity. The ideal candidate will have deep knowledge of transformer models, strong Python/PyTorch skills, and hands-on experience with LLM inference systems.
Step into the Remote Inference Runtime Engineer for LLM & Multimodal AI role at Inferact in United States and grow with us.
Join us at Inferact as our next Remote Inference Runtime Engineer for LLM & Multimodal AI in United States.
We are currently recruiting a Remote Inference Runtime Engineer for LLM & Multimodal AI for our team in United States.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).