×
Register Here to Apply for Jobs or Post Jobs. X

AI​/ML Infrastructure Engineer

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Predii
Full Time position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Salary/Wage Range or Industry Benchmark: 150000 - 210000 USD Yearly USD 150000.00 210000.00 YEAR
Job Description & How to Apply Below

Mid-Level to Senior | Engineering & Platform Ops

US-based - CA preferred, open to West Coast + remote. Hybrid-friendly.

AI/ML Infrastructure Engineer

Mid-Level to Senior | Engineering & Platform Ops

US-based - CA preferred, open to West Coast + remote. Hybrid-friendly.

Address: 2211 Park Blvd, Palo Alto, CA 94306

ABOUT PREDII

Predii builds the intelligence layer that runs the automotive service and parts industry. Our platform, Predii 360, turns messy repair-order, DMS, and parts data into real-time intelligence - powering parts lookup, diagnostics, and repair search for dealership and aftermarket customers at scale, processing billions of repair orders and serving live search at sub-second latency.

We're small, fast, and allergic to red tape. No 12-layer approval chains, no work that disappears into a backlog forever. If you build something here, it ships - and real customers use it. Learn more at

PREDII RESEARCH

We do real research, not just integration. We continue to submit state-of-the-art research on topics including: engineering-diagram and technical-document understanding, domain-calibrated evaluation frameworks for technical content, multi-agent architectures that optimize for correctness and honesty, detecting "confident-but-wrong" failures that standard monitoring misses, moving from reactive detection to causal, explainable prognosis, and multilingual evaluation of technical and repair content. We've found that multi-agent systems that just concatenate outputs get less trustworthy as they get more capable, so we design ours to contest and qualify each other's findings instead.

And we run open-weight models in production at enterprise scale, because repair-grade accuracy shouldn't cost frontier-model money. All of it is deliberately vertical: deep automotive domain expertise applied to automotive problems, not a general-purpose model with an automotive skin.

THE VIBE

We need a AI/ML Infrastructure Engineer who wants more than tickets - someone ready to actually own infrastructure across multiple clouds and help shape how we build. This is real ownership, not busywork. You'll touch:

  • Multi-cloud infra (Azure, GCP, AWS)
  • Kubernetes, CI/CD, automation-everything
  • Security, compliance, access - keeping the house locked
  • Monitoring & reliability - catching problems before customers do
  • Incident response & disaster recovery
  • Cloud cost optimization (yes, we care about the bill too)

Senior folks: expect to shape architecture and mentor the team, not just execute someone else's roadmap.

WHAT YOU'LL ACTUALLY DO
  • Cloud & Platform - build + run deployments on Azure, GCP, AWS with Kubernetes/AKS, Terraform, Ansible, Helm.
  • Dev Ops & CI/CD - ship pipelines that are reliable and repeatable, not held together with duct tape.
  • Reliability & Observability - build monitoring, logging, alerting, tracing; hunt down root causes, not just symptoms.
  • Security & Compliance - RBAC, auth, network security, vuln management, SOC 2 Type 2 controls.
  • Resilience & Ops - own backups, DR, capacity planning, cloud costs, and production support.
  • Keep Leveling Up - evaluate new tools across Dev Ops, infra, and Dev Sec Ops ; you're not stuck with 2019's stack.
  • IT Support - jump in on Windows Server / macOS support when needed.
YOUR TOOLKIT
  • Cloud & Containers: Azure, GCP, AWS, Docker, Kubernetes, AKS
  • Automation: Terraform, Ansible, Helm, Bash, Python
  • CI/CD: Git Hub Actions, Git Lab CI, Azure Dev Ops (or similar)
  • Observability: Grafana, Prometheus, ELK (or equivalent)
  • Networking: TCP/IP, DNS, load balancing, VPNs, firewalls/NSGs, segmentation
  • Identity: RBAC, cloud IAM, Okta, multi-tenant auth
  • Systems: Linux, Windows Server, macOS
  • Security & Compliance: SOC 2 Type 2, access/change controls, business continuity + DR
WHAT YOU BRING
  • 3-5+ years hands-on AI/ML Infrastructure / Dev Ops / SRE / Platform Engineering, in production - not just labs.
  • Strong cloud + Kubernetes chops - Azure/AKS preferred; GCP/AWS/Docker is a big plus.
  • Solid networking and security fundamentals.
  • Comfortable with Terraform, Ansible, Helm, Bash, Python (or similar).
  • Real Git-based CI/CD experience - automated deploys, security scanning included.
  • Battle-tested on observability & prod ops - monitoring,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary