SRE
Listed on 2026-09-08
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AWS
Salary: £68, per year
Requirements:- 3+ years experience in SRE, Platform, or Dev Ops roles within production environments.
- Strong Kubernetes operational experience (on-prem and AWS EKS).
- Hands-on experience defining and operating SLOs/SLIs, alerting, and incident workflows.
- Deep understanding of observability and telemetry (monitoring, logging, tracing).
- Infrastructure as Code with Terraform; experience with Git Ops workflows and CI/CD.
- Scripting proficiency in Python, Bash, or Go.
- Proven ability to balance cost efficiency with reliability and performance.
- Excellent communication skills and the ability to work effectively across multiple teams.
- Experience running chaos engineering experiments is desirable.
- Exposure to high-throughput, low-latency systems is desirable.
- Fin Ops knowledge or cost management practices is desirable.
- AWS certifications (e.g., Solutions Architect, Dev Ops Engineer) are desirable.
- Partner with engineering teams to define, measure, and manage SLOs/SLIs, using error budgets to guide delivery decisions.
- Enhance observability across services (metrics, logs, traces) to detect and resolve issues proactively.
- Lead cost optimisation: monitor spend, right-size workloads, tune autoscaling, and improve infrastructure efficiency.
- Improve production readiness via pre-deployment checks, post-release validation, and robust platform guardrails.
- Introduce and run chaos engineering experiments to strengthen resilience and recovery.
- Automate operational processes to reduce manual intervention and toil across the stack.
- Support major incident response, root-cause analysis, and continual improvement actions.
- Collaborate cross-functionally to raise standards for stability, security, performance, and compliance.
- AWS
- Architect
- Bash
- CI/CD
- Dev Ops
- Git Ops
- Support
- Kubernetes
- Python
- Security
- Terraform
- Cloud
- Incident Management
We are a premier provider of high-volume software solutions for the global iGaming and predictive analytics sector, with a footprint spanning the USA, UK, and Europe. We partner with industry leaders to engineer sophisticated platforms for sports wagering, prize-based systems, and complex market simulation environments. Our vision is to lead the evolution of interactive technology through intelligent, data-driven architecture that ensures seamless user experiences.
We operate with a culture of teamwork, transparency, and technical excellence. This is a full-time permanent Site Reliability Engineer role based in London City on a hybrid basis, with 3 days per week on-site, and a salary of 90,000 per annum.
last updated 36 week of 2026
#J-18808-LjbffrTo Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: