AI Ops Engineer; AI FinOps, Governance, Reliability, and Production Support
Listed on 2026-09-30
-
IT/Tech
SRE/Site Reliability, IT Infrastructure, AI Engineer (Applied/Software), AI Business & Operations
Urgent requirement for AI Ops Engineer (AI Fin Ops, Governance, Reliability, and Production Support) in banking domain is required for our banking clients in Abu Dhabi ,UAE
Design and manage standardized CI/CD pipelines, release workflows, deployment automation, promotion controls, and governance for AI applications, agents, and platform services.
--
Must
Implement operational controls for models and AI assets, ensuring versioning, traceability, compliance, auditability, and safe AI releases across environments.
--Must
Enable canary deployments, controlled rollouts, rollback strategies, AI quality evaluations, monitoring, telemetry, dashboards, runbooks, and production-readiness practices.
--Must
Drive AI Fin Ops, cost visibility, token/model usage monitoring, capacity management, self-service templates, operational playbooks, and enterprise-wide standards for scalable AI operations.
--Must
Banking Domain --Must
Role Purpose
We are seeking an AI Ops Engineer to establish the operational backbone for enterprise AI platforms, enabling application teams to release AI products safely, repeatedly, and role is accountable for production release discipline, LLMOps practices, deployment automation, operational governance, cost visibility, and self-service operating standards for AI-native delivery teams.
- AI Release Engineering & CI/CD: build and design standard release pipelines and promotion controls for AI applications, agents, platform services, and configuration changes across environments, ensuring repeatable deployment, governance, and release evidence.
- LLMOps & AI Lifecycle Management: embed operating controls for models, AI assetsand release for auditable.
- Progressive Delivery:
Implement deployment patterns to reduce production risk, including controlled rollout, canary release, and rollback readiness - Evaluation, Observability & Production Readiness:
Embed AI quality checks, operational telemetry, dashboards, runbooks, and readiness criteria into the delivery lifecycle so AI services are measurable and supportable. - AI Fin Ops & Capacity Governance:
Provide visibility and controls for AI workload consumption, including model usage, token spend, platform capacity, quota management, and optimization opportunities. - Self-Service & Continuous Improvement:
Convert proven operating patterns into reusable templates, release standards, onboarding guidance, operational playbooks, and paved-road workflows that allow teams to move quickly while maintaining enterprise control. - CI/CD & Release Engineering (Git Hub Actions or similar)
- LLMOps / MLOps / AI Lifecycle Management
- Cloud-native Platform Operations & Observability
- AI Fin Ops, Governance, Reliability, and Production Support
- Cross-functional collaboration with Platform, SRE, QA, Security, Architecture, and Product teams.
- Strong production engineering background operating cloud-native, AI, or high-scale API platforms in an enterprise environment.
- Hands-on experience with CI/CD, Git Hub Actions or equivalent automation, and environment management.
- Working knowledge of LLMOps, telemetry, release governance, and production-readiness practices.
- Experience with observability, change control, service reliability, and continuous operational improvement.
Ability to partner with platform engineering, QA/SRE, cybersecurity, architecture, product, and delivery teams to standardize safe and scalable AI operations
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).