Site Reliability Engineer; SRE
Listed on 2026-09-27
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AWS, Systems Engineer
Location
Riyadh, Saudi Arabia
Job Category- Information Technology (IT) & Software
- Engineering & Technical
- Telecommunications
- Others / Miscellaneous
We are seeking an experienced Site Reliability Engineer (SRE) to join a technical team in Riyadh, Saudi Arabia. The role is ideal for professionals with 8+ years of experience in SRE, Dev Ops, production support, or related infrastructure engineering. The successful candidate will work on scalable AWS cloud infrastructure, production reliability, monitoring, automation, and enterprise technology operations.
The SRE will play a key role in improving system reliability and operational efficiency through cloud infrastructure management, observability, incident response, automation, and continuous improvement. The position provides opportunities for career growth, professional development, technical training, and upskilling across AWS, Kubernetes, Dev Ops, cloud security, and modern reliability engineering practices.
Candidates should have strong hands-on experience with AWS Cloud, Kubernetes, Docker, Linux, networking, scripting, monitoring, incident management, and Root Cause Analysis. The ability to join within a maximum of 15 days is required.
Key Responsibilities- Build and manage scalable AWS cloud infrastructure using Terraform and Cloud Formation.
- Implement and maintain monitoring and observability solutions using Prometheus, Grafana, Splunk, or Datadog.
- Manage production incidents and provide effective incident resolution.
- Conduct Root Cause Analysis (RCA) for production issues.
- Define and manage SLOs, SLIs, and error budgets.
- Participate in on-call production support.
- Automate operational activities to improve system reliability and reduce MTTR.
- Manage Kubernetes and Docker environments.
- Support and maintain CI/CD pipelines.
- Administer and troubleshoot Linux environments.
- Apply networking fundamentals to production infrastructure and troubleshooting.
- Develop and maintain operational runbooks.
- Support ITIL and ITSM-based production operations.
- Collaborate with technical teams to improve system availability, reliability, and operational performance.
- Support cloud infrastructure upgrades, automation, and reliability initiatives.
- 8+ years of experience preferred in SRE, Dev Ops, Production Support, or related roles.
- Strong hands-on experience with AWS Cloud
. - Practical experience with Kubernetes and Docker.
- Strong Linux administration and troubleshooting skills.
- Understanding of networking fundamentals.
- Experience with monitoring, observability, incident management, and Root Cause Analysis.
- Proficiency in scripting with Python, Bash, or Go
. - Experience with production operations and reliability engineering practices.
- Knowledge of SLOs, SLIs, error budgets, and MTTR concepts.
- Experience supporting ITIL/ITSM-based environments.
- Terraform or Cloud Formation experience is preferred.
- Jenkins or Git Lab CI/CD experience is preferred.
- Service Now/ITSM exposure is preferred.
- Knowledge of cloud security tools is advantageous.
- Experience working in Cisco or large enterprise environments is preferred.
- Strong analytical, troubleshooting, communication, and problem-solving skills.
- Must be available to join within 15 days maximum
.
The job description does not specify a salary or compensation package.
This position offers opportunities for career growth and professional development within Site Reliability Engineering, Dev Ops, cloud infrastructure, and enterprise production operations. Professionals will gain exposure to AWS, Kubernetes, infrastructure automation, observability, CI/CD, cloud security, incident management, and reliability engineering practices.
The role also provides opportunities to strengthen technical expertise through hands-on experience, continuous upskilling, and industry-relevant cloud and Dev Ops certifications.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).