Member of Technical Staff - ML Performance
Job in
San Francisco, San Francisco County, California, 94102, USA
Listed on 2026-08-05
Listing for:
Modal
Full Time
position Listed on 2026-08-05
Job specializations:
-
Software Development
AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Job Description & How to Apply Below
The Role
We are looking for strong engineers with experience in making ML systems performant you are interested in contributing to open-source projects and Modal's container runtime to push language and diffusion models towards higher throughput and lower latency, we'd love to hear from you!
Requirements- 5+ years of experience writing high-quality, high-performance code.
- Experience working with torch, high-level ML frameworks, and inference engines (vLLM or TensorRT).
- Familiarity with Nvidia GPU architecture and CUDA.
- Experience with ML performance engineering (tell us a story about boosting GPU performance — debugging SM occupancy issues, rewriting an algorithm to be compute-bound, eliminating host overhead, etc).
- Nice-to-have: familiarity with low-level operating system foundations (Linux kernel, file systems, containers, etc).
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×