×
Register Here to Apply for Jobs or Post Jobs. X

Senior Machine Learning Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Acceler8 Talent
Full Time position
Listed on 2026-08-18
Job specializations:
  • Software Development
    AI Engineer (Applied/Software), DevOps, AI Reliability/ Performance Engineer, Machine Learning/ ML Engineer
Salary/Wage Range or Industry Benchmark: 170000 - 240000 USD Yearly USD 170000.00 240000.00 YEAR
Job Description & How to Apply Below

Member of Technical Staff – ML Systems & Inference

San Francisco, CA (Onsite)

I am seeking a Member of Technical Staff to join one of the most exciting AI infrastructure companies building the next generation of inference systems.

As AI models continue to grow in size and complexity, the challenge is no longer simply adding more GPUs, it's about making diverse hardware work together efficiently. This team is building the infrastructure that intelligently executes AI workloads across heterogeneous compute, delivering significant improvements in performance, efficiency, and scalability for production AI applications.

You'll join a small, highly technical engineering team solving some of the hardest problems in AI systems, working across inference runtimes, scheduling, memory management, and distributed infrastructure.

What You'll Do:

  • Design and build production-grade ML inference and model serving systems
  • Optimise latency, throughput, and resource utilisation across large-scale AI workloads
  • Develop execution strategies around batching, scheduling, concurrency, and runtime optimisation
  • Improve KV cache management, memory efficiency, and model execution behaviour
  • Enable new model architectures and inference techniques to run efficiently in production
  • Partner closely with compiler, kernel, networking, and distributed systems engineers to drive end-to-end performance
  • Help shape the architecture of a next-generation AI inference platform powering production workloads at scale

What We're Looking For:

  • Strong software engineering fundamentals
  • Experience building ML inference or model serving systems in production
  • Deep understanding of system performance, memory behaviour, and optimisation under production workloads
  • Experience with inference runtimes such as vLLM
    ,
    TensorRT-LLM
    , or custom serving frameworks is highly desirable
  • Familiarity with batching, scheduling, concurrency, and KV cache management
  • Experience profiling and optimising latency- and throughput-critical systems
  • Strong Python and C++ development experience
  • Comfortable working in a fast-moving, early-stage environment with significant ownership

This is an opportunity to join an exceptionally well-funded AI infrastructure company with a small, world-class engineering team already supporting production deployments for Fortune 500 and AI-native organisations.

You'll work across compiler systems, GPU kernels, distributed scheduling, inference optimisation, and heterogeneous compute, solving difficult engineering challenges that directly impact how modern AI workloads are executed in production.

#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary