Lead Site Reliability Engineer (AWS
Listed on 2026-08-01
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Support
As a Lead Site Reliability Engineer
, you will play a key role in ensuring the reliability, availability, and operational excellence of the British Council's global digital platforms. With a strong specialism in AWS, along with demonstrable experience, you will work closely with engineering teams, architects, senior stakeholders, and our managed service partner, you will design, implement, and continuously improve resilient, scalable, and secure systems. You will leverage expertise in cloud infrastructure, automation, monitoring, performance optimisation, and incident management to deliver high system uptime and support evolving business needs.
the Team
The role sits within the Digital and Technology Engineering team, which partners across the British Council to deliver customer-centric digital products and services. The Engineering function brings together Architecture, Software Engineering, Quality Assurance, and Delivery capabilities to build and maintain world-class digital solutions. Guided by the values of being Open and Committed, Optimistic and Bold, and Expert and Inclusive
, the team champions collaboration, innovation, digital inclusion, and engineering best practices to create reliable, high-performing technology that enhances the experience of customers and colleagues worldwide.
You will contribute to shaping and delivering the British Council's Site Reliability / Dev Ops strategy by implementing best practices that improve the availability, resilience, security, and performance of our digital platforms. You will use your expertise in cloud infrastructure, automation, monitoring, Fin Ops and performance optimisation to identify opportunities for improvement, resolve complex reliability challenges, and ensure systems remain scalable and highly available.
You will also leverage system metrics, user feedback, and data analytics to drive continuous improvement while promoting inclusive, user-centric digital experiences.
You will build strong cross-functional relationships, advocate for Site Reliability Engineering across the organisation, and collaborate with teams to define performance expectations and deliver robust solutions. In addition, you will mentor fellow engineers to build up their AWS expertise, have involvement with Azure responsibilities where required, contribute to the Site Reliability Engineering community of practice, share knowledge and best practices, and foster a culture of collaboration, continuous learning, and operational excellence.
Rolespecific skillsTechnical
- Cloud Expertise:
Expertise with major cloud platforms, specifically AWS, but ideally also Microsoft Azure. - Infrastructure Automation:
Excellent working knowledge using tools like Terraform or Ansible to automate infrastructure tasks. - Containerisation and Orchestration:
Knowledge of Docker and Kubernetes for application deployment and management. - Monitoring and Observability:
Experience with tools such as Azure Monitor, ELK stack, or similar to monitor system health, analyse logs, and troubleshoot production issues. - Programming ,Scripting and Infrastructure as Code:
Ability to write and maintain scripts in languages like Python or Typescript for automation and Terraform for IaC - Networking:
Understanding of load balancers, DNS setups, and basic network security practices. - Database Management:
Experience in managing and optimising relational and No
SQL databases. - SLIs, SLOs, SLAs:
Familiarity with setting and monitoring service level metrics to maintain system reliability.
- Problem-solving:
Ability to quickly diagnose and resolve system-related issues. - Communication:
Articulating technical information effectively to varied audiences. - Collaboration:
Working closely with developers, other SREs, and technical teams to ensure system robustness
- Operational Management:
Overseeing daily operations, ensuring processes are adhered to, and optimizing workflows. - Mentoring and sharing skills and expertise with the wider SRE team.
- Teamwork:
Actively contributing to team meetings, discussions, and brainstorming sessions, sharing knowledge and best practices. - Chaos Engineering:
Some exposure…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: