Site Reliability Engineer; SRE UAE
Listed on 2026-09-29
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, Cybersecurity
Position Summary
Innovations Global is urgently seeking an experienced, proactive Site Reliability Engineer (SRE) to champion the continuous availability, scalability, and performance of mission-critical banking and financial technology platforms in Abu Dhabi, United Arab Emirates. Operating in a full-time onsite capacity for a 1-year renewable contract, you will bridge the divide between software engineering and systems operations to ensure enterprise infrastructure resilience. With a minimum of five years of hands-on experience, you will spearhead observability, drive end-to-end automation, lead rapid incident recovery, and enforce stringent Site Reliability Engineering principles across hybrid cloud environments.
This position provides an exceptional opportunity for a technical engineering professional to make a transformative impact on premier banking infrastructure in the UAE.
Job Description
As a Site Reliability Engineer (SRE) at Innovations Global supporting a premier enterprise banking environment in Abu Dhabi, you will hold operational accountability for the health, availability, performance, and efficiency of high-throughput transactional applications and distributed systems. You will collaborate closely with cross-functional software engineering teams, Dev Ops squads, database administrators, and cyber security teams to establish resilient deployment pipelines and maintain robust production ecosystems.
Your core technical mandate involves architecting and administering enterprise Linux and Unix servers, orchestrating microservices utilizing Docker and Kubernetes, and managing scalable workloads across leading cloud platforms (AWS, Azure, or GCP). You will implement proactive monitoring and observability frameworks, define and track Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Service Level Agreements (SLAs), and eliminate operational toil through Python and Bash automation scripting.
Additionally, you will direct incident response triage, execute root cause analysis (RCA), optimize CI/CD release workflows, and troubleshoot complex TCP/IP enterprise networking bottlenecks. This role requires rigorous diagnostic discipline, deep systems acumen, and the capability to maintain zero-downtime reliability within a fast-paced financial services environment.
- Ensure the maximum reliability, availability, performance, and operational efficiency of enterprise banking applications and underlying cloud infrastructure.
- Implement, tune, and manage full-stack monitoring, telemetry, and observability platforms to capture proactive operational insights and real-time alerts.
- Eliminate operational toil by designing, building, and maintaining automated workflows and operational scripts using Python, Bash, or Shell scripting.
- Lead rapid incident management, triage system outages, conduct detailed root cause analysis (RCA), and implement permanent corrective remediations.
- Deploy, configure, manage, and scale containerized application workloads utilizing Kubernetes clusters and Docker environments.
- Administer, optimize, and maintain high-performance enterprise Linux and Unix server operating systems in Tier-compliant hosting environments.
- Establish, monitor, and report on core reliability engineering metrics, including Service Level Agreements (SLAs), Service Level Objectives (SLOs), and Error Budgets.
- Optimize and support automated CI/CD deployment pipelines, ensuring secure, reliable, and frictionless software releases into production.
- Minimum 5+ years of dedicated professional experience in Site Reliability Engineering (SRE), Dev Ops, or Linux/Unix Systems Engineering.
- Strong technical expertise in Linux and Unix…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).