Machine Learning Operations; MLOps Engineer UAE
Listed on 2026-07-10
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Machine Learning/ ML Engineer, AI Engineer (Applied/Software)
Machine Learning Operations (MLOps) Engineer
Location:
Abu Dhabi, United Arab Emirates
Job Type: Full Time
Position SummaryIT People Gulf is seeking a skilled MLOps Engineer to manage and scale production AI and machine learning environments. This role ensures reliable model deployment, automated ML pipelines, platform stability, and operational excellence across enterprise AI initiatives.
DetailedJob Description
As an MLOps Engineer, you will build, automate, and maintain the infrastructure that powers AI and machine learning solutions in production. You will work closely with data scientists, AI engineers, and platform teams to deploy models, optimize inference performance, and manage the model lifecycle. The role requires strong expertise in containerization, orchestration, cloud infrastructure, CI/ CD pipelines, and observability tools. You will scale AI workloads, support GPU-based environments, and implement best practices for monitoring, security, and reliability.
Key Responsibilities- Deploy, manage, and monitor machine learning models in production environments
- Design and maintain end‑to‑end MLOps pipelines for model training, testing, deployment, and monitoring
- Implement and optimize CI/CD processes for AI and machine learning workloads
- Manage containerized applications using Docker and Kubernetes
- Support model serving environments and API‑based AI services
- Monitor model performance, system health, and resource utilization
- Troubleshoot infrastructure, deployment, and application performance issues
- Manage GPU‑enabled environments for AI and deep learning workloads
- Implement logging, observability, and monitoring frameworks for production systems
- Support model versioning, upgrades, rollback strategies, and controlled releases
- Collaborate with AI, engineering, and operations teams to improve platform scalability and reliability
- Ensure platform security, governance, and operational best practices are maintained
- 3–5 years of experience in MLOps, Dev Ops, Cloud Engineering, AI Infrastructure, or Platform Engineering
- Strong hands‑on experience with Docker and Kubernetes
- Experience building and managing CI/CD pipelines
- Proficiency in Python for automation and infrastructure tasks
- Experience deploying and managing machine learning models in production
- Knowledge of model serving architectures and API‑based deployments
- Experience with monitoring, logging, and observability platforms
- Understanding of GPU infrastructure and AI workload optimization
- Experience working with cloud or private cloud environments
- Strong troubleshooting, automation, and problem‑solving skills
- Excellent communication and collaboration abilities
- Experience with MLflow, Kubeflow, Airflow, or similar MLOps platforms
- Familiarity with AWS, Azure, Google Cloud, or Open Shift environments
- Knowledge of Large Language Models (LLMs) and Generative AI deployments
- Infrastructure as Code (Terraform, Ansible) experience
- Experience with data pipelines, feature stores, and model governance frameworks
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).