Site Reliability Engineer; Global
Listed on 2026-08-25
-
IT/Tech
SRE/Site Reliability, IT Support
Job Family:
Operations | Function:
Information Systems / Computer Operations |
Location:
Durban | Salary:
Rneg up to R45k p/m (depending on experience)
We're looking for an Engineer:
Global Operations to join our 24/7 production operations team, supporting global iGaming platforms. In this role, you'll ensure system stability, reliability, and availability through incident management, root cause analysis, and technical problem-solving — balancing reactive incident response with proactive improvement. You'll play a key part in leading post-incident reviews, supporting operational readiness assessments, validating AI-assisted operational decisions, and mentoring junior colleagues to build team capability.
You'll Do
Participate in 24/7 shift rotations, managing incident response and triage — owning alert acknowledgement, prioritisation, and investigation decisions, resolving issues within established procedures, and escalating to specialists only when needed. Conduct root cause analysis (RCA) investigations, determine technical solutions, collaborate with engineering teams on complex problems, and take accountability for resolution outcomes. Coordinate and facilitate severity-tiered post-incident reviews, extracting lessons learned and driving organisational learning.
Own and maintain incident knowledge and documentation — knowledge base structure, standards, and accessibility for the team. Mentor and coach junior team members, transferring technical knowledge and developing team proficiency in investigation and response procedures. Identify operational bottlenecks and inefficiencies, propose and implement workflow refinements, reduce alert noise, and optimise incident response procedures. Support AI-assisted operational decision validation — establishing validation processes, escalation criteria, curating incident data for agent training, and optimising human-in-the-loop model performance.
Assess operational readiness for new features and system changes, including SLI/SLO implications and operational risk identification. Evaluate and recommend operational tools and platforms, assessing supportability, reliability, and integration options. Drive automation and toil-reduction initiatives — identifying repetitive manual tasks and implementing efficiency improvements.
Education Advanced Diploma or Bachelor's Degree in Information Technology, Computer Science, Computer Engineering, Information Systems, Cybersecurity or an equivalent technical field. Experience 2–4 years of progressive experience in IT operations, technical support, cybersecurity or systems administration, with demonstrated capability in incident investigation, troubleshooting, and process improvement. Skills Incident Management, Troubleshooting, Root Cause Analysis, System Administration, Technical Documentation, Problem-Solving (Developing–Intermediate) Communication Skills (Intermediate–Advanced) Process Improvement, Knowledge Management, Escalation Management (Developing–Intermediate) Automation, Agile Methodology, Quality Assurance, Change Management, Influencing Skills, ITIL 4 (Awareness–Developing) Knowledge Broad knowledge across incident management, post-incident learning, process improvement, emerging operational technologies, and operational readiness assessment Advanced understanding of incident response practices — alert triage, investigation methodology, root cause analysis, escalation procedures, and documentation standards Demonstrated technical knowledge of production systems, application architecture, and infrastructure components Developing proficiency with AI-assisted operational tools and automation opportunities Working knowledge of operational readiness assessment, SLOs, and quality standards
Why Join UsBe part of a global operations team supporting platforms that operate across multiple countries and continents. Work at the intersection of traditional operations and AI-assisted automation, helping shape how human-in-the-loop processes evolve. Grow your technical and leadership skills through mentoring, cross-functional collaboration with engineering teams, and exposure to a broad range of operational challenges. Contribute directly to platform reliability and customer trust in a regulated, high-availability environment.
This is a full-time role requiring participation in a 24/7 shift rotation.
#J-18808-LjbffrTo Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: