×
Register Here to Apply for Jobs or Post Jobs. X

Lead Site Reliability Engineer

Job in Nottingham, Nottinghamshire, NG1, England, UK
Listing for: London Stock Exchange Group
Full Time position
Listed on 2026-07-29
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Support
Salary/Wage Range or Industry Benchmark: 90000 - 120000 GBP Yearly GBP 90000.00 120000.00 YEAR
Job Description & How to Apply Below

Role Profile We are seeking an accomplished and forward-thinking technical leader to join our Risk Intelligence organisation as the Tech Lead 2d Site Reliability Engineering (SRE) for the Risk Screening Product Line. Based in Nottingham, UK, this role will lead the reliability engineering function for critical Screening applications supporting the EMEA region. The successful candidate will be accountable for ensuring the availability, scalability, performance, and security of business-critical platforms while driving operational resilience and service excellence across a complex, cloud-native technology landscape.

As a key technical leader, you will combine deep expertise in Site Reliability Engineering, cloud operations, and modern software delivery practices with a passion for building high-performing engineering teams. You will lead reliability strategies, incident management, automation initiatives, and continuous improvement efforts, while partnering closely with product, engineering, and business stakeholders to align technology outcomes with organisational objectives. This role offers an exciting opportunity to influence engineering culture, mentor and develop talent, and establish best-in-class operational practices within a growing and innovative UK-based technology organisation.

Core Responsibilities Technical Ownership for Identity & Fraud Platforms Lead 24×7 reliability operations for Screening services, ensuring uptime, performance, and compliance across WC1 and World Check verify Applications. Define and implement SLOs, SLIs, and error budgets; maintain reliability scorecards and drive improvements through engineering backlogs. Act as Major Incident Commander during critical outages; lead blameless post-incident reviews and ensure learnings are institutionalized.

Drive automation-first mindset across incident response, deployment, compliance, and observability. Collaborate with product and engineering teams to embed non-functional requirements and secure-by-design principles into delivery pipelines. Co-own cloud reliability roadmap with platform teams; standardize tooling for observability, ITSM, and incident communication. Ensure DR readiness, runbook quality, and resilience patterns are consistently applied across services. People Leadership Lead and mentor a team of SRE engineers, fostering a culture of ownership, learning, and engineering excellence.

Drive career development, performance management, and technical capability growth across the team. Collaborate with HR and Talent teams to build local hiring pipelines and support workforce planning. Promote well-being and psychological safety within the team; ensure compliance with health and safety standards. Represent Nottingham site in global SRE forums; contribute to offshore strategy and location planning. Partner with vendors and staffing partners to manage workforce augmentation and ensure delivery quality.

Support BCP/DR planning and ensure site-level operational readiness for critical events.

Required Skills and Experience 8+ years in production operations, SRE, or Dev Ops roles, with at least 3+ years in a people management capacity. Proven experience managing cloud-native services on Azure and AWS, including:
Azure SQL, Cosmos DB, Application Gateway, Key Vault, Storage, DNS, Load Balancer, Virtual Machines, Azure Machine Learning, Sentinel AWS Lambda, ECS, RDS, Cloud Watch Strong understanding of SRE principles (SLOs, SLIs, error budgets, incident response). Hands-on experience with container orchestration (Kubernetes, Docker). Proficiency in CI/CD pipelines and infrastructure-as-code tools (Terraform, Git Hub Actions, Jenkins). Familiarity with observability platforms (Datadog, Big Panda, Open Telemetry).

Experience working with identity platforms and/or fraud detection systems. Excellent communication and stakeholder management skills; ability to influence across technical and business domains. Strong analytical mindset with a focus on measurable outcomes and continuous improvement.

Personal Attributes Empathetic and inclusive leader, committed to team well-being and growth Strategic thinker with a bias for…

Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary