AI Ops Engineer | IDC Technologies
Listed on 2026-10-03
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Infrastructure
Position Summary
IDC Technologies is seeking a seasoned and hands‑on AI Ops Engineer to join our technology team in Abu Dhabi. In this pivotal role, you will bridge the gap between traditional Dev Ops, Site Reliability Engineering (SRE), and cutting‑edge artificial intelligence infrastructure. You will be responsible for orchestrating, scaling, and monitoring large language model operations (LLMOps), managing containerized Kubernetes clusters, and driving cloud‑native automation across multi‑cloud environments.
This opportunity requires a senior technical professional with 7+ years of experience who can ensure high availability, optimal performance, and cost‑effective resource utilization for enterprise AI workloads.
Job Description
As the AI Ops Engineer, you will take full ownership of designing, deploying, and maintaining the operational pipelines that power our enterprise AI and generative AI platforms. Your day‑to‑day responsibilities include managing Kubernetes clusters, implementing robust CI/CD deployment pipelines, configuring infrastructure‑as‑code using Terraform, and deploying comprehensive observability and monitoring frameworks. You will closely monitor AI Fin Ops to optimize cloud spending across Azure and AWS, ensuring cost efficiency without compromising model throughput or reliability.
Furthermore, you will collaborate with data science and software engineering teams to streamline model containerization, automate incident response, and establish resilient production workflows.
- Design, deploy, and manage scalable infrastructure supporting AI, Machine Learning, and LLMOps workloads in enterprise environments.
- Administer and optimize container orchestration platforms using Kubernetes, Docker, and container registries.
- Implement and maintain robust CI/CD pipelines for seamless model deployment, testing, and continuous integration.
- Automate cloud infrastructure provisioning, configuration management, and network security policies using Terraform.
- Establish comprehensive observability, logging, and monitoring solutions to track system health, model latency, and error rates.
- Drive AI Fin Ops initiatives to monitor, analyze, and optimize cloud infrastructure costs across Azure and AWS platforms.
- Collaborate with cross‑functional data science and software engineering teams to troubleshoot production issues and scale AI pipelines.
- Ensure high system availability, disaster recovery readiness, and adherence to enterprise security and compliance standards.
- Minimum of 7 years of professional experience in Dev Ops, Site Reliability Engineering (SRE), or Cloud Infrastructure Engineering.
- Proven hands‑on experience implementing AI/LLMOps pipelines, managing large language model deployments, and serving infrastructure.
- Extensive technical proficiency in Kubernetes administration, container orchestration, and microservices architecture.
- Strong multi‑cloud deployment and administration experience across Microsoft Azure or Amazon Web Services (AWS).
- Demonstrated expertise in infrastructure‑as‑code tools, specifically Terraform and configuration management frameworks.
- Hands‑on experience building automated CI/CD pipelines and integrating observability tools (e.g., Prometheus, Grafana, Datadog).
- Familiarity with AI Fin Ops practices, cloud cost management, and resource allocation optimization.
- Bachelor’s degree in Computer Science, Information Technology, Software Engineering, or a related technical discipline.
- Professional certifications such as Certified Kubernetes Administrator (CKA) or AWS/Azure Dev Ops Engineer Expert.
- Prior experience working with vector databases (e.g., Pinecone, Milvus, Qdrant) and LLM serving…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).