×
Register Here to Apply for Jobs or Post Jobs. X

Machine Learning Engineer, LLM Inference Optimization in Sonoma

Job in Sonoma, Sonoma County, California, 95476, USA
Listing for: NLP PEOPLE
Full Time position
Listed on 2026-07-07
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer, AI Reliability/ Performance Engineer, Backend Developer
Salary/Wage Range or Industry Benchmark: 120000 - 160000 USD Yearly USD 120000.00 160000.00 YEAR
Job Description & How to Apply Below

Job Description

About Us

GMI Cloud is a fast‑growing AI infrastructure company backed by Headline VC and one of only seven cloud providers worldwide to earn NVIDIA’s prestigious Reference Platform Cloud Partner designation. We operate 8 of our own GPU clusters across the U.S. and Asia, delivering a full spectrum of services from GPU compute to AI model inference API solutions. As an NVIDIA Reference Platform Cloud Partner, our infrastructure meets the highest standards for performance, security, and scalability in AI deployments.

We empower AI startups and enterprises to “build AI without limits,” providing everything they need to prototype, train, and deploy AI models quickly and reliably.

About this role

GMI Cloud is building the leading inference optimization solution and the most advanced token platform in the global token market — and we are hiring world‑class Machine Learning Engineers to make GMI the new industry benchmark for LLM serving performance, cost efficiency, and production reliability.

This role is for engineers who want to live at the frontier of LLM inference systems. You will drive the research, validation, and productionization of the most advanced inference optimization techniques, and turn them into real competitive advantage over top open‑source baselines (vLLM, SGLang, and so on). Our charter is not just to adopt what’s published — it is to define the recipes, ship the optimizations, and contribute back to the community that the rest of the industry follows.

You will focus on B200‑first optimization, with support for H200 evolution, across core domains including quantization, speculative decoding, KV cache & memory management, prefill/decode disaggregation, and system‑level inference optimization. You will work closely with platform and infrastructure teams to transform cutting‑edge ideas into measurable gains in latency, throughput, cost efficiency, and production scalability.

Key Responsibilities
  • Drive frontier research and engineering in LLM inference optimization across one of the four focus tracks (Speculative Decoding, Quantization, PD Disaggregation, KV Cache & Memory) while contributing across the full optimization stack.
  • Develop next‑optimization strategies for large‑scale LLM serving across model execution, runtime systems, and production inference platforms — with B200 as the primary target and H200 as a continuing platform.
  • Advance state‑of‑the‑art techniques in quantization (NVFP4 / MXFP4 / FP8, QAT), speculative decoding (EAGLE‑3, MTP, DFlash, Model Opt, Spec Forge), KV cache & memory management (LMCache / HiCache / NV KVBM, paged attention, prefix‑aware routing), and PD disaggregation (NVIDIA Dynamo, KV‑aware router/planner, fault recovery).
  • Drive system‑level optimization across scheduling, batching, routing, gateway orchestration, adapter serving, and end‑to‑end inference efficiency.
  • Build scalable optimization frameworks, performance methodologies, and benchmark infrastructure that allow GMI to stay ahead of the industry as models, hardware, and serving patterns evolve.
  • Productionize cutting‑edge ideas into real customer workloads — measured by TTFT, ITL, throughput, goodput, tail latency, quality, and unit token cost.
  • Engage with and contribute to the open‑source community (vLLM, SGLang, TensorRT‑LLM, NVIDIA Dynamo / Model Opt, Flash Infer, LMCache, etc.) — read upstream code, file issues, send PRs, and publish tech blogs and case studies.
  • Collaborate closely with platform, infrastructure, and product teams to make inference optimization a core technical advantage of GMI Cloud.
Required Skills
  • Strong hands‑on experience with LLM inference systems and performance optimization on modern GPUs.
  • Solid understanding of inference metrics and tradeoffs, including TTFT, ITL, throughput, goodput, tail latency, GPU utilization, memory efficiency, and quality/cost tradeoffs.
  • Experience with one or more modern serving stacks such as SGLang, vLLM, TensorRT‑LLM, NVIDIA Dynamo, or Triton.
  • Deep familiarity with GPU‑based inference, model serving architecture, and production bottlenecks around compute, memory bandwidth, KV‑cache behavior, and scheduling.
  • Demonstrable depth in at…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary