LLM Inference Performance Engineer - GPU Kernel Optimizer
Job in
Folsom, Sacramento County, California, 95630, USA
Listed on 2026-08-12
Listing for:
Intel
Full Time
position Listed on 2026-08-12
Job specializations:
-
Software Development
AI Engineer (Applied/Software), AI Reliability/ Performance Engineer
Job Description & How to Apply Below
Intel is seeking a performance-focused AI Infrastructure Engineer to push LLM inference to its limits on Intel GPUs. You will optimize end-to-end inference pipelines, profile cross-stack bottlenecks, and develop high-performance kernels for attention, MoE, and quantization.
You will also upstream improvements into vLLM, SGLang, and PyTorch while shaping future GPU roadmaps. Work spans from kernel development to open-source contributions, with a hybrid work model and a competitive total
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×