×
Register Here to Apply for Jobs or Post Jobs. X

ML Engineer, Inference Platform

Job in Sunnyvale, Santa Clara County, California, 94087, USA
Listing for: Israelvcforum
Full Time position
Listed on 2026-07-04
Job specializations:
  • Software Development
    Machine Learning/ ML Engineer, AI Engineer (Applied/Software), AI Reliability/ Performance Engineer, DevOps
Salary/Wage Range or Industry Benchmark: 195000 - 298000 USD Yearly USD 195000.00 298000.00 YEAR
Job Description & How to Apply Below
Position: Staff ML Engineer, Inference Platform

Job Description

Hybrid. This role is categorized as hybrid, meaning the successful candidate is expected to report to the Sunnyvale Technical Center, CA at least three times per week, or another frequency dictated by the business. This job is eligible for relocation assistance.

About the Team

The ML Inference Platform is part of the AI Compute Platforms organization within Infrastructure Platforms. Our team owns the cloud-agnostic, reliable, and cost‑efficient platform that powers GM’s AI efforts, serving teams developing autonomous vehicles (L3/L4/L5) and other AI‑driven products. We enable rapid innovation by optimizing for high‑priority, ML‑centric use cases. Our platform supports the serving of state‑of‑the‑art machine learning models for experimental and bulk inference, focusing on performance, availability, concurrency, and scalability.

We are committed to maximizing GPU utilization across platforms (B200, H100, A100, etc.) while maintaining reliability and cost efficiency.

About the Role

We are seeking a Staff ML Infrastructure Engineer to build and scale robust compute platforms for ML workflows. In this role, you’ll work closely with ML engineers and researchers to ensure efficient model serving and inference in production for workflows such as data mining, labeling, model distillation, simulations, and more. This is a high‑impact opportunity to influence the future of AI infrastructure  will shape the architecture, roadmap, and user experience of a robust ML inference service supporting real‑time, batch, and experimental inference needs.

The ideal candidate brings experience designing distributed systems for ML, strong problem‑solving skills, and a product mindset focused on platform usability and reliability.

What you’ll be doing
  • Design and implement core platform backend software components.
  • Collaborate with ML engineers and researchers to understand critical workflows, parse them into platform requirements, and deliver incremental value.
  • Lead technical decision‑making on model serving strategies, orchestration, caching, model versioning, and auto‑scaling mechanisms.
  • Drive the development of monitoring, observability, and metrics to ensure reliability, performance, and resource optimization of inference services.
  • Proactively research and integrate state‑of‑the‑art model serving frameworks, hardware accelerators, and distributed computing techniques.
  • Lead large‑scale technical initiatives across GM’s ML ecosystem.
  • Raise the engineering bar through technical leadership, establishing best practices.
  • Contribute to open‑source projects and represent GM in relevant communities.
Minimum Requirements
  • 8+ years of industry experience, focused on machine learning systems or high‑performance backend services.
  • Expertise in one or more relevant coding languages:
    Go, Python, C++.
  • Expertise in ML inference and model serving frameworks (e.g., Triton, Ray Serve, vLLM).
  • Strong communication skills and a proven ability to drive cross‑functional initiatives.
  • Experience working with cloud platforms such as GCP, Azure, or AWS.
  • Ability to thrive in a dynamic, multi‑tasking environment with evolving priorities.
Preferred Qualifications
  • Hands‑on experience building ML infrastructure platforms for model serving/inference.
  • Experience designing interfaces, APIs, and clients for ML workflows.
  • Experience with the Ray framework, and/or vLLM.
  • Experience with distributed systems and large‑scale data processing.
  • Familiarity with telemetry and feedback loops to inform product improvements.
  • Familiarity with GPU hardware acceleration and optimizations for inference workloads.
  • Contributions to open‑source ML serving frameworks.
Why Join Us?

If you’re excited to tackle complex engineering challenges, see the impact of your work in real‑world autonomous vehicle applications, and help shape the future of AI infrastructure at GM, this is the team for you.

Compensation

The compensation information is a good faith estimate based on applicable state laws. The expected base compensation for this role is $195,000 - $298,000. Actual base compensation may vary based on factors relevant to the position.

Bonus potential:
An incentive pay program offers…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary