Site Reliability Engineer; SRE DevOps & Cloud Engineering UAE
Abu Dhabi, UAE/Dubai
Listed on 2026-10-03
-
IT/Tech
Cloud Computing: Infrastructure & Operations, AWS, SRE/Site Reliability, Systems Engineer
Location
Abu Dhabi, United Arab Emirates — Local Candidates Only
Job Category- Information Technology (IT) & Software
- Engineering & Technical
- Telecommunications
- Remote / Work-from-Home Opportunities
An opportunity is available for an experienced Site Reliability Engineer (SRE) / Production Engineer in Abu Dhabi, UAE. The role focuses on owning and supporting customer-facing production systems with strong hands-on expertise in AWS, Amazon EKS, Kubernetes, Python, automation, and cloud infrastructure.
The successful candidate will be responsible for production reliability, cloud infrastructure, observability, incident management, automation, and continuous improvement. Strong experience with Kubernetes scaling, AWS services, Infrastructure as Code, CI/CD, and production troubleshooting is required. The position offers opportunities for professional development, cloud and Dev Ops training, relevant certification, and long-term career growth in site reliability and cloud engineering.
Key Responsibilities- Own and support customer-facing production systems.
- Develop Python-based automation and engineering solutions.
- Manage AWS and Amazon EKS/Kubernetes production environments.
- Configure and optimize HPA, Karpenter/Cluster Autoscaler, networking, ingress, storage, and RBAC.
- Perform Kubernetes cluster upgrades and production maintenance.
- Manage AWS Lambda and serverless architectures.
- Administer IAM/IRSA, VPC, S3, ECR, API Gateway, SQS, SNS, Event Bridge, and managed databases.
- Manage Helm and Kustomize deployments.
- Implement Git Ops and Infrastructure as Code using Terraform, CDK, or Cloud Formation.
- Support CI/CD pipelines and deployment automation.
- Monitor production systems using Cloud Watch, X-Ray, and Cloud Trail.
- Participate in incident management, on-call support, postmortems, and production troubleshooting.
- Define and monitor SLOs and reliability objectives.
- Troubleshoot Linux, containers, networking, and distributed systems.
- Work with stakeholders and technical teams to improve system reliability and performance.
- 5+ years of experience in SRE, Dev Ops, Production Engineering, or Software Engineering.
- Strong Python development and automation experience.
- Deep hands-on AWS and Amazon EKS/Kubernetes production experience.
- Strong knowledge of HPA, Karpenter/Cluster Autoscaler, networking, ingress, storage, RBAC, and Kubernetes upgrades.
- Production experience with AWS Lambda and serverless architectures.
- Strong knowledge of IAM/IRSA, VPC, S3, ECR, API Gateway, SQS/SNS/Event Bridge, and managed databases.
- Experience with Helm, Kustomize, Git Ops, Terraform, CDK, Cloud Formation, and CI/CD.
- Strong AWS observability experience with Cloud Watch, X-Ray, and Cloud Trail.
- Experience with incident management, on-call operations, postmortems, SLOs, and production troubleshooting.
- Strong Linux, container, networking, and distributed-systems fundamentals.
- Excellent communication and stakeholder-management skills.
- Bachelor’s degree in Computer Science, Engineering, or a related field.
- Candidates must be currently based locally in Abu Dhabi/UAE.
Salary and compensation details were not provided by the employer and therefore are not listed.
The role offers opportunities for:
- Professional development in SRE, Dev Ops, and cloud engineering.
- Advanced AWS and Kubernetes technical exposure.
- Cloud infrastructure and automation training.
- Relevant AWS, Kubernetes, Dev Ops, and cloud certification opportunities.
- Experience managing large-scale production environments.
- Career progression in Site Reliability Engineering, Dev Ops, and Cloud Engineering.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).