ML Ops Engineer
Listed on 2026-08-30
-
IT/Tech
AWS, SRE/Site Reliability, Data Engineering, Cloud Computing: Infrastructure & Operations
Job Description
You'll design, deploy, and maintain ML pipelines on AWS and Databricks, automating the full model lifecycle from training to deployment and monitoring. You'll implement CI/CD workflows for ML systems, ensure model reproducibility and versioning, optimize compute costs and performance, and collaborate closely with data scientists and engineers to bring models reliably into production.
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances.
If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy:
- 3–5 years of experience in MLOps, data engineering, or a related field
- Strong Python skills, including writing production-grade, testable code
- Hands-on experience with AWS services (e.g., S3, Sage Maker, Lambda, ECS/EKS, IAM)
- Solid working knowledge of Databricks (jobs, workflows, MLflow, Unity Catalog)
- Experience with CI/CD tooling (e.g., Git Hub Actions, Git Lab CI) and infrastructure-as-code (e.g., Terraform)
- Familiarity with containerization (Docker) and orchestration concepts
- Understanding of ML lifecycle management: experiment tracking, model registries, monitoring, and retraining
- High level of comfort with Linux
- Experience with SQL
- Strong problem-solving and analytical skills
- Excellent communication and collaboration skills
- Experience with Spark/PySpark at scale
- Knowledge of data governance and security best practices
- Relevant AWS or Databricks certifications
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).