Senior MLOps Engineer
Listed on 2026-07-23
-
Software Development
DevOps, Machine Learning/ ML Engineer, Cloud Engineer - Software
About the Role
As the Senior MLOps Engineer I, you will help turn the models built by our ML Scientists, Data Scientists, and Perception Engineers into reliable, production‑grade services. You will work on the infrastructure, pipelines, and tooling that take a model or an LLM/agent‑backed workflow from a research notebook to a fully monitored deployment across multiple industry verticals, including our model registry, deployment pipelines, and the cloud infrastructure our AI/ML platform depends on.
This role sits at the intersection of R&D, Software Engineering, and Dev Ops. You will work daily with our R&D team to understand what a model needs to run in production (compute, data inputs, versioning, post‑processing), and partner closely with the Platform and Dev Ops teams to provision necessary infrastructure, permissions, and deployment pathways. You will also contribute to broader automation initiatives, providing deployment visibility and pipeline reliability that let R&D, Software, Product, and Ops teams move in lockstep.
The day‑to‑day will include maintaining and extending our model registry, building and debugging deployment pipelines and cloud infrastructure, and setting up model and pipeline monitoring and testing. You will troubleshoot issues such as failed deployments, permissions errors, or inconsistent environments, and help shape and document standards for how models move from staging to production. Most importantly, you will serve as a key communicator ensuring R&D goals and challenges are well understood by Software Engineering and Dev Ops teams.
Responsibilities- Partner with Scientists: Work directly and iteratively with ML Scientists, Data Scientists, and Perception Engineers to translate experimental, research‑oriented code into dependable, scalable production services without slowing down their research velocity.
- Cross‑Functional
Collaboration:
Coordinate with Dev Ops and Software Engineering teams on infrastructure requests and shared data pipeline needs, and support broader automation initiatives and team goals. - Model Registry, Deployment & Release Management: Maintain and improve model registry and deployment pipelines, and help implement safer release practices (e.g., shadow deployments, rollback procedures) to reduce risk.
- Cloud Infrastructure & CI/CD: Build, maintain, and troubleshoot cloud infrastructure and CI/CD pipelines that ML workloads run on, working closely with Engineering and Dev Ops teams on shared tooling, infrastructure‑as‑code, and cost optimization for compute‑heavy workloads.
- Monitoring, Drift & Reproducibility: Implement monitoring and observability for models and pipelines in production, help R&D track model performance and drift over time, and support experiment tracking and dataset/model versioning.
- Ongoing Maintenance & Platform Support: Keep deployed ML systems healthy over time with dependency and infrastructure upgrades, capacity and cost management, data pipeline upkeep, and retraining or redeployment support, and extend support as needs evolve.
- Standards & Documentation: Help define and document conventions for model versioning, deployment promotion, and model documentation/lineage, and build tools to allow scientists and engineers to self‑serve.
- Bachelor’s degree in Computer Science, Software Engineering, Data Engineering, or a related field; typically 4+ years of professional experience in MLOps, ML platform engineering, or infrastructure engineering supporting machine learning teams.
- Solid, applied knowledge of MLOps practices, with the ability to work independently across varied production scenarios and
** escalate
* * only genuinely complex or ambiguous problems. - Demonstrated experience working directly with researchers or ML scientists. You understand research workflows and can translate them into reliable services and productionized models without becoming a bottleneck. You serve as a key link, communicating R&D goals and challenges to Software Engineering and Dev Ops teams.
- Strong Python skills and solid software engineering fundamentals (testing, code review, version control).
- Hands‑on experience with a major cloud platform (e.g.,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).