×
Register Here to Apply for Jobs or Post Jobs. X

Member of Technical Staff - ML Performance

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Modal
Full Time position
Listed on 2026-08-03
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 180000 - 260000 USD Yearly USD 180000.00 260000.00 YEAR
Job Description & How to Apply Below

About Us:

AI needs a new infrastructure layer. We’re building it at Modal.

Every era of computing brought new workloads that previous infrastructure couldn’t support: mainframes, databases, and the cloud. Each time, the company that rebuilt the layer underneath defined the decade. AI is no different, except it touches everything instead of one slice, and the window to build the layer underneath it is open right now.

Our customers include category-defining companies like Lovable, Ramp, Cognition, Door Dash, and Suno. They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it’s simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.

We recently raised a $355M Series C at a $4.65B valuation, led by General Catalyst and Redpoint Ventures. We’ve crossed $300M+ ARR and grown fivefold since September.

Our team includes creators of popular open-source projects (e.g.,Seaborn,Luig i), academic researchers, international olympiad medalists, and experienced engineering and product leaders with decades of experience.

The Role:

We are looking for strong engineers with experience in making ML systems performant  you are interested in contributing to open-source projects and Modal’s container runtime to push language and diffusion models towards higher throughput and lower latency, we’d love to hear from you!

Requirements:
  • 5+ years of experience writing high-quality, high-performance code.
  • Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).
  • Familiarity with Nvidia GPU architecture and CUDA.
  • Experience with ML performance engineering (tell us a story about boosting GPU performance — debugging SM occupancy issues, rewriting an algorithm to be compute-bound, eliminating host overhead, etc).
  • Nice-to-have: familiarity with low-level operating system foundations (Linux kernel, file systems, containers, etc).
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary