×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Cardiff, Cardiff City Area, CF10, Wales, UK
Listing for: Camwebdir
Full Time position
Listed on 2026-09-02
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below

Darktrace is a global leader in AI for cybersecurity that keeps organizations ahead of the changing threat landscape every day. Founded in 2013, Darktrace provides the essential cybersecurity platform protecting nearly 10,000 organizations from unknown threats using its proprietary AI.

The Darktrace Active AI Security Platform™ delivers a proactive approach to cyber resilience to secure the business across the entire digital estate – from network to cloud to email. Breakthrough innovations from our R&D teams have resulted in over 200 patent applications filed. Darktrace’s platform and services are supported by over 2,400 employees around the world. To learn more, visit

Job Description About the Role

We’re looking for a Site Reliability Engineer (SRE) to bring deep expertise in a key reliability domain and help shape the future of our platform reliability strategy.

SRE sits at the heart of our operational trifecta alongside Platform Engineering and Dev Sec Ops . In this role, you’ll act as the go‑to authority in your area of specialism, working across teams to embed best practices, solve complex reliability challenges, and improve system resilience at scale.

Unlike a generalist SRE, this role focuses on a core domain of expertise—such as observability, performance engineering, data infrastructure reliability, security‑focused SRE, or network reliability—while influencing reliability standards across the wider engineering organisation.

Key Responsibilities Domain Expertise & Strategy
  • Act as the subject matter expert in your chosen reliability domain
  • Define and implement standards, frameworks, and best practices across SRE, Platform Engineering, and Dev Sec Ops
  • Stay current with industry trends and bring innovative ideas into the organisation
Engineering & Delivery
  • Design and implement solutions to complex, cross‑cutting reliability challenges
  • Build tooling, automation, and frameworks to improve system resilience and scalability
  • Lead deep‑diving investigations into systemic issues and drive long‑term fixes
Collaboration & Platform Integration
  • Partner with Platform Engineering to ensure your domain is embedded within the internal developer platform
  • Collaborate with Dev Sec Ops  to integrate security, compliance, and resilience practices
  • Contribute to cross‑team initiatives that improve reliability across the stack
Incident & Operational Excellence
  • Play a key role in incident response, particularly within your specialism
  • Contribute to on‑call rotations and continuous improvement of operational processes
  • Develop runbooks, documentation, and training materials to support teams
What You’ll Bring Essential
  • Proven experience in Site Reliability Engineering, Dev Ops, or infrastructure engineering
  • Deep expertise in at least one of the following areas:
    • Observability & monitoring (metrics, logging, distributed tracing)
    • Performance engineering & capacity planning
    • Data infrastructure reliability (databases, streaming, pipelines)
    • Security‑focused SRE (hardening, compliance automation, secrets management)
    • Network reliability & traffic management
  • Strong programming skills (e.g. Go, Python, or similar)
  • Experience with cloud platforms (AWS, GCP, Azure) and Kubernetes
  • Strong communication skills, with the ability to explain complex technical concepts clearly
  • Self‑driven with the ability to identify and prioritise high‑impact work independently
Desirable
  • Experience building internal developer platforms or tooling
  • Contributions to open‑source, technical blogs, or public speaking
  • Experience working in regulated environments
  • Familiarity with SLO frameworks and error budget management
  • Relevant certifications in your specialist domain
Success Measures
  • Improved reliability and performance within your domain of specialism
  • Adoption of best practices across SRE, Platform Engineering, and Dev Sec Ops
  • Reduction in incidents and faster resolution times
  • Scalable, well‑integrated solutions within the internal platform
  • Strong collaboration across teams and measurable improvements in operational maturity
Why Join Us?
  • Shape reliability strategy in a modern, cloud‑native engineering environment
  • Work on complex, high‑impact systems at scale
  • Collaborate with expert teams…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary