×
Register Here to Apply for Jobs or Post Jobs. X

Sr MLOps Engineer

Job in Sunnyvale, Santa Clara County, California, 94087, USA
Listing for: Intuitive
Full Time position
Listed on 2026-08-11
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 160000 - 271000 USD Yearly USD 160000.00 271000.00 YEAR
Job Description & How to Apply Below

Company Description

It started with a simple idea: what if surgery could be less invasive and recovery less painful? Nearly 30 years later, that question still fuels everything we do  a global leader in robotic-assisted surgery and minimally invasive care, our technologies—like the da Vinci surgical system and Ion—have transformed how care is delivered for millions of patients worldwide.

We’re a team of engineers, clinicians, and innovators united by one purpose: to make surgery smarter, safer, and more human. Every day, our work helps care teams perform with greater precision and patients recover faster, improving outcomes around the world.

The problems we solve demand creativity, rigor, and collaboration. The work is challenging, but deeply meaningful—because every improvement we make has the potential to change a life.

If you’re ready to contribute to something bigger than yourself and help transform the future of healthcare, you’ll find your purpose here.

Job Description Primary Function of Position

In this role, you will be responsible for designing, building, and maintaining the infrastructure and tools necessary to support the entire machine learning lifecycle, from development to deployment. You will work closely with ML engineers and software developers across Intuitive to ensure that machine learning models are seamlessly integrated into our systems and deliver value  ideal candidate is an independent and fast-paced engineer with excellent problem-solving skills and practical working knowledge of modern ML development techniques.

Essential

Job Duties
  • Bootstrap and maintain a production-grade Kubernetes cluster, including CNI networking and storage integration
  • Deploy and configure ML orchestration tooling (e.g., Metaflow) and artifact/dataset storage solutions to support reproducible ML workflows
  • Validate GPU node health and configuration across heterogeneous hardware (B200, L40S, A6000, V100), including driver/CUDA standardization and topology checks
  • Design and execute team migration playbooks, working directly with engineering teams to port workflows, migrate datasets/artifacts, and roll out tool updates
  • Write and maintain runbooks, architecture documentation, and disaster recovery procedures
  • Participate in on-call rotation and incident response for platform-level issues
  • Collaborate with IT/Security on identity integration, access control, and compliance requirements
  • Continuously evaluate and adopt infrastructure best practices for reliability, cost, and developer experience
Required Skills and Experience
  • 3+ years of experience in infrastructure, Dev Ops, or MLOps roles, or equivalent practical experience
  • Demonstrated experience operating Kubernetes in production (networking, storage, RBAC, troubleshooting)
  • Strong scripting/automation skills in Python and/or Bash; comfort with Infrastructure-as-Code tools (Ansible, Helm, Terraform, or similar)
  • Hands-on experience with at least one distributed storage system (S3, MinIO, Net App, or similar)
  • Experience building or maintaining CI/CD pipelines (Git Lab CI, ArgoCD, or equivalent)
  • Solid understanding of Linux systems administration and networking fundamentals
  • Excellent communication and documentation skills, with the ability to write clear runbooks and migration guides
  • High degree of autonomy and comfort working across the full stack, iteratively building solutions
Required Education And Training
  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field; or equivalent experience
Preferred Skills And Experience
  • Experience with ML orchestration frameworks (Metaflow, MLflow, Kubeflow, or similar)
  • Familiarity with GPU infrastructure (NVIDIA drivers, CUDA, NVLink/NUMA topology, MIG partitioning)
  • Prior experience in a regulated industry (healthcare, finance, or similar) where auditability and access control are critical
  • Experience leading or supporting large-scale infrastructure migrations with multiple stakeholder teams
Additional Information

Due to the nature of our business and the role, please note that Intuitive and/or your customer(s) may require that you show current proof of vaccination against certain diseases including COVID-19. Details…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary