More jobs:
Senior Site Reliability Engineer
Job in
Ruddington, Nottingham, Nottinghamshire, NG1, England, UK
Listed on 2026-07-10
Listing for:
Experian Ltd
Full Time
position Listed on 2026-07-10
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support, Systems Engineer
Job Description & How to Apply Below
Salary
Salary: £55,000 - £95,000 per year
Requirements- Deep expertise with AWS services
- Advanced knowledge of monitoring and observability tools
- Strong leadership skills with the ability to set clear direction, align team efforts with organizational goals, and maintain high motivation and engagement
- Excellent communication skills with the ability to explain complex ideas, solutions, and feedback to technical and non-technical stakeholders
- Ability to manage conflict constructively and facilitate consensus
- Proven track record of building secure, mission-critical, high-volume transaction web-based software systems, preferably in regulated environments such as finance and insurance
- Hands-on technologist with experience in software development and leading an SRE team
- Define and implement SRE best practices across the organization
- Lead production support, engineering, disaster recovery, automation, and cloud operations
- Mentor and guide a team of SREs and support their growth
- Collaborate with senior stakeholders to align reliability goals with business objectives
- Establish SLIs, SLOs, and SLAs for critical services and ensure adherence
- Drive initiatives to improve system resilience and reduce operational toil
- Design systems that detect and remediate issues without manual intervention
- Develop self-healing systems and runbook automation
- Use tools such as Gremlin, Chaos Monkey, and AWS FIS to simulate outages and improve fault tolerance
- Act as the primary escalation point for critical production issues
- Lead major incident response, root cause analysis, and postmortems
- Perform detailed post-incident investigations to identify underlying causes
- Document findings and share learnings to prevent recurrence
- Implement preventive measures and continuous improvement processes
- Champion monitoring, logging, and alerting strategies using tools such as Prometheus, Grafana, ELK, and AWS Cloud Watch
- Build real-time dashboards to visualize system health and reliability metrics
- Configure intelligent alerting based on anomaly detection and thresholds
- Combine metrics, logs, and traces to support root cause analysis and reduce MTTR
- Apply AIOps or ML-based anomaly detection for proactive reliability management
- Work closely with development teams to integrate reliability into application design and deployment
- Promote a culture of shared responsibility for uptime and performance across engineering teams
- AWS
- Cloud
- Cloud Watch
- ELK
- Grafana
- Support
- Marketing
- Prometheus
- Web
- Dev Ops
Position Requirements
10+ Years
work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×