Senior Lead Site Reliability Engineer
Listed on 2026-09-06
-
IT/Tech
SRE/Site Reliability, Systems Engineer
At the heart of everything we do is our vision to change lives every day and our mission to grow The National Lottery responsibly and champion its impact.
We are Allwyn UK part of the Allwyn Entertainment Group a multi-national lottery operator with a market-leading presence acrossthe USA (Michigan and Illinois) and Europe including
Czech Republic Austria Greece Cyprusand Italy.
While the main contribution of The National Lottery to society is through the funds togood causes at Allwynwe put our purpose and values at the heart of everything we us as we embark on a once-in-a-lifetime large scale transformation journey by creating a National Lottery that delivers more money togood causes.
Welltalk a bit more about us further down the page but for now letstalk about the role and whowerelooking for
A bit about the role
At Allwyn the Senior/Lead Site Reliability Engineer is responsible for technical leadership of reliability engineering across the digital estate ensuring high availability performance and resilience of customer-facing systems during both normal operation and peak lottery events.
The role combines hands-on engineering incident leadership and ownership of the SRE improvement backlog and reporting working across platform product and operational teams.
Objectives of the role
- Own reliability outcomes across services using SLOs SLIs and error budgets
- Improve availability latency and scalability across Instant-Win and Draw-based platforms
- Lead incident response and operational readiness including peak jackpot events
- Drive automation and platform maturity reducing manual operational effort
- Establish clear reporting on reliability incidents and service health trends
- Supporting with thought-leadership and developing long-term roadmap
Whatyoullbe doing
Reliability engineering & technical leadership
- Define and govern SLOs / SLIs / error budgets across critical services
- Lead reliability design reviews across:
- Web & mobile platforms
- Player Identity & Protection systems
- CMS and Geolocation services
- Drive architecture improvements for resilience (failover degradation scaling patterns)
Production operations & incident leadership
- Act as incident commander for major incidents and high-severity events
- Lead 1-in-4 on-call rotation covering:
- Out-of-hours incident diagnosis
- Peak jackpot proactive monitoring
- Own end-to-end incident lifecycle:
- Detection triage resolution post-incident review
- Ensure blameless post-mortems with clear remediation ownership
Observability & service insight
- Define and evolve observability strategy using:
- Splunk (log analytics)
- Cloud Watch (AWS telemetry)
- Grafana (metrics visualisation)
- Quantum Metric (user behaviour insight)
- Standardise:
- Alerting quality and signal-to-noise ratio
- Dashboards aligned to SLOs and customer impact
- Drive correlation between technical signals and user experience
Automation & platform engineering
- Lead automation of operational processes using Terraform and scripting
- Improve deployment and release safety (CI/CD progressive delivery patterns)
- Key contributor to transition strategy for ECS EKS (Kubernetes adoption)
- Reduce operational toil through tooling self-healing observability and platform improvements
- Empower Level-1 operational teams with safe controlled access to the tools they need to operate autonomously
Capacity & performance engineering
- Own capacity planning for:
- High-concurrency draw events
- Traffic spikes during jackpots
- Lead performance optimisation:
- Latency reduction
- Throughput scaling
- Cost efficiency (AWS utilisation and associated log costs observability license consumption)
Backlog ownership & reporting
- Own and prioritise the SRE backlog balancing:
- Reliability improvements
- Technical debt
- Automation opportunities to reduce/offload toil
- Produce structured reporting covering:
- SLO performance
- Incident trends and MTTR
- Platform health and risk areas
- Provide clear updates to engineering leadership and business stakeholders
Collaboration & culture
- Embed SRE practices across engineering teams
- Mentor engineers and SREs on:
- Reliability engineering
- Observability
- Incident management
- Promote a culture of automation measurement and continuous improvement
- Supporting with thought-leadership and developing…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: