×
Register Here to Apply for Jobs or Post Jobs. X

Lead Site Reliability Engineer (AWS

Job in Cardiff, Cardiff City Area, CF10, Wales, UK
Listing for: British Council
Full Time position
Listed on 2026-08-01
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Support
Salary/Wage Range or Industry Benchmark: 90000 - 140000 GBP Yearly GBP 90000.00 140000.00 YEAR
Job Description & How to Apply Below
Position: Lead Site Reliability Engineer (AWS)

As a Lead Site Reliability Engineer
, you will play a key role in ensuring the reliability, availability, and operational excellence of the British Council's global digital platforms. With a strong specialism in AWS, along with demonstrable experience, you will work closely with engineering teams, architects, senior stakeholders, and our managed service partner, you will design, implement, and continuously improve resilient, scalable, and secure systems. You will leverage expertise in cloud infrastructure, automation, monitoring, performance optimisation, and incident management to deliver high system uptime and support evolving business needs.

About

the Team

The role sits within the Digital and Technology Engineering team, which partners across the British Council to deliver customer-centric digital products and services. The Engineering function brings together Architecture, Software Engineering, Quality Assurance, and Delivery capabilities to build and maintain world-class digital solutions. Guided by the values of being Open and Committed, Optimistic and Bold, and Expert and Inclusive
, the team champions collaboration, innovation, digital inclusion, and engineering best practices to create reliable, high-performing technology that enhances the experience of customers and colleagues worldwide.

Main responsibilities

You will contribute to shaping and delivering the British Council's Site Reliability / Dev Ops strategy by implementing best practices that improve the availability, resilience, security, and performance of our digital platforms. You will use your expertise in cloud infrastructure, automation, monitoring, Fin Ops and performance optimisation to identify opportunities for improvement, resolve complex reliability challenges, and ensure systems remain scalable and highly available.

You will also leverage system metrics, user feedback, and data analytics to drive continuous improvement while promoting inclusive, user-centric digital experiences.

You will build strong cross-functional relationships, advocate for Site Reliability Engineering across the organisation, and collaborate with teams to define performance expectations and deliver robust solutions. In addition, you will mentor fellow engineers to build up their AWS expertise, have involvement with Azure responsibilities where required, contribute to the Site Reliability Engineering community of practice, share knowledge and best practices, and foster a culture of collaboration, continuous learning, and operational excellence.

Role

specific skillsTechnical
  • Cloud Expertise:
    Expertise with major cloud platforms, specifically AWS, but ideally also Microsoft Azure.
  • Infrastructure Automation:
    Excellent working knowledge using tools like Terraform or Ansible to automate infrastructure tasks.
  • Containerisation and Orchestration:
    Knowledge of Docker and Kubernetes for application deployment and management.
  • Monitoring and Observability:
    Experience with tools such as Azure Monitor, ELK stack, or similar to monitor system health, analyse logs, and troubleshoot production issues.
  • Programming ,Scripting and Infrastructure as Code:
    Ability to write and maintain scripts in languages like Python or Typescript for automation and Terraform for IaC
  • Networking:
    Understanding of load balancers, DNS setups, and basic network security practices.
  • Database Management:
    Experience in managing and optimising relational and No

    SQL databases.
  • SLIs, SLOs, SLAs:
    Familiarity with setting and monitoring service level metrics to maintain system reliability.
Soft Skills
  • Problem-solving:
    Ability to quickly diagnose and resolve system-related issues.
  • Communication:
    Articulating technical information effectively to varied audiences.
  • Collaboration:

    Working closely with developers, other SREs, and technical teams to ensure system robustness
Leadership
  • Operational Management:
    Overseeing daily operations, ensuring processes are adhered to, and optimizing workflows.
  • Mentoring and sharing skills and expertise with the wider SRE team.
  • Teamwork:
    Actively contributing to team meetings, discussions, and brainstorming sessions, sharing knowledge and best practices.
  • Chaos Engineering:
    Some exposure…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary