×
Register Here to Apply for Jobs or Post Jobs. X

Expert Site Reliability Engineer

Job in Riyadh, Riyadh Region, Saudi Arabia
Listing for: TAWANTECH
Full Time position
Listed on 2026-09-03
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 240000 - 320000 SAR Yearly SAR 240000.00 320000.00 YEAR
Job Description & How to Apply Below

Purpose:

To drive the reliability, availability, scalability, and operational resilience of critical technology services by applying advanced software engineering, automation, observability, and reliability engineering practices.

Main

Duties and Responsibilities:

  • Define and implement advanced reliability engineering practices across critical technology services.
  • Establish and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability targets.
  • Design automation to reduce manual operational activities and improve system resilience.
  • Develop and enhance monitoring, observability, alerting, and incident detection capabilities.
  • Lead technical analysis and resolution of complex production incidents.
  • Conduct root-cause analysis and drive permanent corrective and preventive actions.
  • Design solutions to improve system availability, scalability, capacity, and disaster resilience.
  • Identify reliability risks and recommend architectural and engineering improvements.
  • Drive performance engineering and capacity planning for critical services.
  • Provide advanced technical guidance and mentorship on SRE practices.
  • Promote automation and engineering approaches that reduce operational toil and improve service reliability.
QUALIFICATIONS & REQUIREMENTS
  • Bachelor's degree in Computer Science, Software Engineering, IT, or a related field
    .
  • 5+ years of experience in Site Reliability Engineering, Dev Ops, Platform Engineering, or related roles.
  • Strong experience in cloud platforms, Kubernetes, and production environments
    .
  • Strong knowledge of monitoring, observability, alerting, SLIs, SLOs, and reliability metrics
    .
  • Hands-on experience with automation, scripting, CI/CD, and Infrastructure as Code (Terraform).
  • Proven experience in complex incident management, troubleshooting, and Root Cause Analysis (RCA).
  • Strong understanding of high availability, scalability, performance engineering, capacity planning, and disaster recovery
    .
  • Experience driving reliability improvements and reducing operational toil through automation
    .
  • Strong analytical, problem-solving, and technical leadership skills.
  • Experience in Banking, Fin Tech, or Payment environments is preferred.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary