ML Infrastructure Engineer
Listed on 2026-08-30
-
Software Development
Machine Learning/ ML Engineer, Cloud Engineer - Software, AI Engineer (Applied/Software), DevOps
About the Role
This is a hands-on ML Infrastructure Engineer role at an early-stage enterprise AI startup, where you'll own the end-to-end inference and model-serving infrastructure that keeps production AI agents running reliably and 'll sit at the intersection of ML and platform engineering, directly shaping the systems that power real-world, high-stakes deployments in regulated industries like insurance, banking, and healthcare.
What You'll DoOwn inference and model-serving infrastructure end to end, from design through production deployment.
Build and scale systems that enable AI agents to run reliably and efficiently under increasing concurrency.
Collaborate closely with ML and infrastructure teams to ensure seamless integration and performance optimization.
Identify infrastructure bottlenecks and drive cross-functional solutions across engineering teams.
5+ years of experience building and operating ML inference systems, model-serving platforms, or ML infrastructure in production.
Hands-on experience designing and scaling inference-serving infrastructure using frameworks such as Tensor Flow Serving, Torch Serve, Triton, KServe, or equivalent custom systems.
Strong track record optimizing production ML systems for latency, throughput, and reliability at scale.
Experience with containerization and orchestration (Docker, Kubernetes) for deploying and scaling ML workloads.
Experience building distributed systems that handle concurrent requests and manage resource allocation under load.
Proficiency with observability and debugging tooling for production systems (e.g., Prometheus, Grafana, ELK, distributed tracing).
Cloud platform experience on AWS, GCP, or Azure for deploying and managing ML systems.
Proficiency in at least one systems or backend language — Python, Go, Rust, C++, or Java.
Nice to have: experience with knowledge graphs, semantic search, or graph databases (e.g., Neo4j, Amazon Neptune); real-time or low-latency inference systems; agentic or multi-step reasoning pipelines; enterprise data infrastructure or integration platforms.
On-site in San Mateo, CA. No visa sponsorship is available for this role.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).