SRE Manager
Listed on 2026-08-30
-
IT/Tech
SRE/Site Reliability
The Team Open Betis a global leader in betting and gaming entertainment, trusted by over 200 partners to create memorable winning moments for millions of players worldwide. From processing bets during iconic events like theFIFA World Cupand Super Bowlto pioneering next-gen products like
Bet Builder, we continuously redefine the player experience with high-quality content,cutting-edge technology, and advanced player protection tools.
For over 25 years, our unbeatable platform has powered the most recognizable betting brands, ensuring peak performance with
100% uptime, unmatched scale, and speed. With 85 licenses, 20 World Lottery Association operators on our customer roster, and a team of 1,200+ experts across 14 countries, weremainat the heart of the industry.
We're looking for an SRE Manager to lead our Site Reliability Engineering team, with a particular focus on observability and performance engineering. You'll build and lead the team responsible for how we see, measure, and understand the health of our platform — setting the standards for monitoring, alerting, and performance testing that keep our systems running reliably. You'll balance hands-on technical leadership with people management, working closely with Dev Ops, Infrastructure, and Database Engineering to raise the bar on reliability across the estate.
What you’ll be doing- Lead and grow the SRE team, providing day-to-day people management, coaching, and career development for a group of reliability and performance engineers
- Own the observability strategy across the platform — metrics, logging, tracing, and alerting — equipping Dev Ops, Infrastructure, and Engineering teams with self-service dashboards and insight rather than gatekeeping the data
- Own the evaluation, budget, and vendor relationship for observability tooling, balancing capability, cost, and operational fit
- Explore and adopt AI-assisted observability capabilities — anomaly detection, predictive alerting, and AI-driven root-cause analysis — to help the team spot issues earlier and resolve them faster
- Own the performance engineering practice, including load testing, capacity planning, and performance benchmarking
- Define and drive SLIs, SLOs, and error budgets for critical platform services, working with Engineering to embed reliability targets into delivery
- Act as an escalation point for major incidents, contributing root cause analysis and ensuring learnings feed back into monitoring, alerting, and performance improvements
- Partner with Dev Ops, Infrastructure, and Database Engineering to close observability gaps across the hybrid on-premises and cloud estate
- Ensure the team maintains clear runbooks and playbooks for common failure scenarios, making reliability knowledge repeatable rather than dependent on individual expertise
- Champion a proactive reliability culture, shifting the team from reactive firefighting toward prevention through better tooling, automation, and standards
- Proven experience leading a Site Reliability Engineering or Performance Engineering team, including direct people management responsibility
- Deep hands-on background in observability — metrics, logging, tracing, and alerting
- Strong experience with performance engineering practices, including load testing, capacity planning, and performance benchmarking
- A track record of defining and ope rationalising SLIs, SLOs, and error budgets in a production environment
- Experience acting as an escalation point for major incident response, contributing to root cause analysis through to concrete reliability improvements
- Comfort working across hybrid on-premises and cloud environments, ideally within a high-availability, transaction-heavy domain
- Hands-on AWS experience, with a focus on cost optimization and leveraging AI-powered tools to drive efficiency across cloud infrastructure
- Working knowledge of Kubernetes and Rancher, enough to operate confidently across our hybrid on-premises and cloud platforms
- An AI-native approach to observability, with exposure to AI or LLM-assisted tools for anomaly detection and root-cause analysis a plus
- Strong communication skills, with the ability to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).