×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff - Inference Research

Job in San Francisco, San Francisco County, California, 94102, USA
Listing for: Modal
Full Time position
Listed on 2026-07-08
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Job Description & How to Apply Below

The Role

Most of the value of owning a model shows up at serving time. We're building a platform that covers the whole life of an LLM -- train it, deploy it, observe it -- and inference is where teams feel the difference every day. We already run elastic inference, sandboxes, distributed volumes, and multi-node training, and we control the infrastructure underneath, so the serving stack is ours to shape rather than something we resell.

What you'll do:

  • Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spiky serverless traffic, and whatever else the research agenda calls for.

  • Train custom speculators against real production traffic and feed what you learn back into target models -- acceptance length is the metric that decides the win.

  • Work directly with customers alongside our Forward Deployed Engineers to deploy and tune models, and bring what you learn back into the research.

  • Carry and expand collaborations with outside research labs, for example:

    • our work with ZLab on DFlash, a speculator design built on KV injection and blockwise parallel drafting

    • our work with SGLang on specdec and multimodal inference performance

    • our work on Flash Attention 4 kernels

  • Work with engineering to turn frontier serving techniques into products: primitives for disaggregation, fast weight refresh for models that keep training after deployment, observability for quality and latency in production, or even a next-generation inference engine.

  • Help shape the research agenda. None of the above is prescriptive; your work will help guide our future.

Requirements
  • A research-leaning or systems background in LLM inference, with work you can point to.

  • Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling.

  • A record of shipping research or systems that other people build on, whether in a lab or in industry.

  • The drive to independently take a research bet from idea to result, working in the open with the rest of the team.

  • Ability to work in-person, in our NYC or San Francisco office.

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary