More jobs:
Senior Machine Learning Ops Engineer
Job in
San Francisco, San Francisco County, California, 94199, USA
Listed on 2026-07-21
Listing for:
Jobtailor
Full Time
position Listed on 2026-07-21
Job specializations:
-
Software Development
Machine Learning/ ML Engineer, DevOps, Cloud Engineer - Software, AI Engineer (Applied/Software)
Job Description & How to Apply Below
Responsibilities
- Architect, design, deploy, and operate scalable cloud-based MLOps platforms and workflows that enable efficient training, evaluation, deployment, monitoring, and lifecycle management of AI/ML models.
- Own the technical strategy and evolution of ML infrastructure, identifying architectural bottlenecks and driving cross‑functional initiatives to improve developer productivity, experimentation velocity, scalability, and operational efficiency.
- Build robust, reliable, and automated systems that enable teams to ship new models and features rapidly while maintaining high standards for quality, reproducibility, observability, security, and production reliability.
- Define and implement infrastructure optimization strategies that balance performance, scalability, reliability, and cost across cloud and compute resources.
- Evaluate emerging tools, technologies, and industry best practices in MLOps, cloud infrastructure, and ML systems, and lead their adoption where they can meaningfully improve ML development and production workflows.
- Establish engineering best practices for ML infrastructure, including system design, code quality, testing, CI/CD, monitoring, documentation, and operational readiness.
- Provide technical leadership and mentorship to engineers, lead design and code reviews, and help raise the engineering quality and technical capabilities of the broader team.
- Partner closely with ML engineers, researchers, data engineers, and product teams to translate evolving AI/ML requirements into scalable and maintainable infrastructure solutions.
- Drive complex, ambiguous infrastructure projects from technical strategy and architecture through implementation, production deployment, and long‑term operational ownership.
- A Bachelors Degree or a Masters Degree in Computer Science, Electrical Engineering, or a related field.
- Core
Skills:
General Software Engineering skills with 6+ years of programming experience in Python and the surrounding tooling ecosystem along with familiarity in Linux and expertise in infrastructure, cloud and/or MLOps. - Personal Attributes:
Team player, good communication skills, self‑starter. - Strong teamwork and communication skills to collaborate with cross‑functional teams, including ML and software engineers.
- Nice to Have:
Experience building MLOps pipelines for deep learning based perception solutions on AWS or GCP.
Demonstrates expertise in architecting and deploying scalable cloud-based MLOps platforms, with a strong focus on optimizing ML infrastructure for performance, reliability, and cost. Proven ability to lead cross‑functional teams and mentor engineers while implementing best practices in software engineering and ML workflows.
#J-18808-LjbffrPosition Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×