Lead Machine Learning Operations Engineer
Listed on 2026-07-24
-
IT/Tech
Machine Learning/ ML Engineer, SRE/Site Reliability
Lead Machine Learning Operations Engineer
Location:
Burbank, CA 91505.
Employment type:
Full‑Time, On‑Site.
We’re hiring a Lead Machine Learning Operations Engineer to own the operational excellence, observability, reliability, and governance layer around our personalization and recommendation ML systems.
Our recommendation models retrain and deploy frequently. You will define how we detect model behavior changes, diagnose issues quickly, and prevent bad deployments from reaching customers.
This is a lead‑level IC role: you’ll set technical direction and drive adoption across ML Engineering, Dev Ops, Platform Engineering, Data Engineering, and Product.
Sitting within ML Platform and Infrastructure, you’ll partner closely with ML engineers who own model development. You’re not expected to build infrastructure from scratch, but you’ll define what good looks like, evaluate tooling, and own the day‑to‑day operational layer.
Responsibilities- Own the ML production reliability strategy
- Define and lead the operational strategy for production ML systems, including monitoring, traceability, deployment safety, incident response, and post‑deployment validation.
- Set the standards ML teams use to assess model health, performance, and trustworthiness in production.
- Own model traceability and governance, ensuring clear lineage and driving adoption of model registry and metadata tooling across ML teams.
- Build end‑to‑end ML observability across the full ML signal path: data arrival, feature freshness, distribution stability, candidate generation, ranking behavior, model metrics, serving latency, and SLA performance.
- Partner with stakeholders to define post‑deployment metrics covering model quality, system reliability, business guardrails, and degradation indicators.
- Detect data drift, feature drift, model behavior changes, and silent failures proactively through thresholding, alerting, anomaly detection, and release‑over‑release monitoring.
- Lead diagnostic tooling and root‑cause analysis, building dashboards, logs, and workflows that progress quickly from “recommendations look off” to root cause.
- Own ML deployment safety, defining automated gates that prevent bad models or data from being promoted to production, and establishing validation checks, rollback criteria, canary strategies, shadow testing, and release health reviews.
- Own incident response practices for ML systems, including rollback playbooks, hotfix strategies, severity definitions, tradeoff frameworks, communications, and post‑mortems.
- Partner with Dev Ops, Data, and ML Engineering to embed operational requirements into development and deployment workflows.
- Establish reusable patterns, playbooks, and standards, and mentor engineers on reliability, observability, and operational rigor.
- 5+ years of experience in machine learning engineering, ML platform, applied ML, MLOps, data platform, reliability engineering, or a related technical role.
- Demonstrated experience operating production ML systems, including monitoring, deployment, incident response, model validation, data quality, or reliability ownership.
- Experience leading technical initiatives across multiple engineering teams, especially where success required influencing architecture, tooling, standards, or adoption.
- Hands‑on experience with model registries, feature stores, ML metadata systems, production monitoring, model deployment pipelines, or ML observability platforms.
- Solid knowledge of end‑to‑end ML systems, including training data, features, model artifacts, offline validation, online serving, post‑deployment metrics, and business outcome measurement.
- Ability to reason about ML operational failure modes: stale features, distribution shift, training‑serving skew, delayed labels, and offline‑online metric gaps.
- Solid SQL skills and comfort investigating data quality, feature distributions, model outputs, pipeline behavior, and production anomalies.
- Track record of cross‑functional collaboration with Platform, Data, and ML Engineering to deliver production‑grade operational capabilities.
- Strong written and verbal communication skills, including the ability to explain ML system health,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).