Site Reliability Engineer
Listed on 2026-09-03
-
Software Development
AWS
Site Reliability Engineer
Home-based (Remote-first with occasional travel to our Hemel Hempstead office and off-site events)
Permanent Full Time
Competitive salary bonus and benefits
About the role
Are you a hands-on Site Reliability Engineer who thrives at the intersection of cloud code and reliability If so we want to hear from you! Haven is looking for a technically strong SRE to join our Product Technology function and help raise the bar for how our digital platforms perform scale and recover.
Youll work shoulder to shoulder with engineering teams and Product Tech Leads to design implement and support the systems that guests owners and colleagues rely on every day. From CI/CD pipelines and observability through to database reliability incident management and disaster recovery this role touches every layer of our stack.
This is also a great time to join. AI-assisted and agentic engineering tooling is a core part of how we build and youll play an active part in shaping how those tools are adopted safely and effectively across the team. If you enjoy reducing toil driving up standards and engineering forwards through uncertainty this one is for you.
Your Opportunity
- Support infrastructure engineering from design through to implementation solving complex challenges around developer productivity and delivery velocity
- Design and improve CI/CD processes and pipelines using Git Actions shaping our release process and deployment strategies so we consistently deliver high-quality releases.
- Partner with developers and engineers to troubleshoot build and deployment issues and unblock delivery
- Contribute to and maintain internally developed engineering tools
- Own monitoring tracing and observability so we are first to know when something is (or is about to be) an issue and can diagnose it quickly
- Drive database reliability across our relational and non-relational estate (RDS PostgreSQL/Aurora PostgreSQL Oracle MS SQL Open Search) including critical legacy systems
- Champion security by ensuring the right technical and business controls are in place including supporting accreditation processes
- Proactively identify performance scalability and load improvements that reduce response times
- Define our approach to incident management including runbooks and response practices
- Ensure disaster recovery documentation controls and training are in place and ready when needed
- Shape our release process and deployment strategies so CI/CD consistently delivers high quality releases
- Partner with our Platform Team to define guardrails and cloud policies
- Share cloud best practice across engineering
- Actively use AI-assisted and agentic coding tools as part of your day to day work and help the team get the most out of them
What wed like you to bring
- Deep knowledge of AWS and a strong grasp of cloud engineering and architectural best practice
- Proven experience with CI/CD and configuration management tools technologies and practices (ideally Git Actions)
- Hands-on experience with Infrastructure as Code ideally Terraform
- Experience with containerisation using Docker and orchestration with Kubernetes
- Working knowledge of NodeJS and Type Script
- Understanding of observability and monitoring incident management disaster recovery and security practice
- Solid grounding in cloud and network security principles and technologies (TLS OAuth2 Active Directory and similar)
- Experience supporting and maintaining critical legacy database systems (Oracle MS SQL) alongside modern engines (PostgreSQL/Aurora PostgreSQL Open Search)
- Active day-to-day use of AI-assisted and agentic coding tools (for example coding harnesses Claude Code) as a core part of your engineering practice
- Understanding of AI agent architectures and how agentic tools apply to software development
- A consultative advisory style with the ability to coach and share expertise on complex topics such as cloud architecture and AI/agentic system adoption
- Comfortable working with uncertainty and engineering forwards
- AWS or other cloud accreditations (desirable not essential)
- Exposure to Azure or GCP alongside AWS (desirable)
Why Haven
Haven is the UKs leading holiday operator welcoming around 4 million guests every year…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: