Site Reliability Engineer
Job in
London, Greater London, W1B, England, UK
Listed on 2026-09-09
Listing for:
Trainline
Full Time
position Listed on 2026-09-09
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Our job is to make the green travel choice the best choice. By building a better train travel experience, we help more people choose rail - creating a positive impact for customers, our business and the planet.
We’re a team of more than 1,000 Train liners from over 50 nationalities, working across London, Paris, Barcelona, Milan, Edinburgh and Madrid. Now is a brilliant time to join us and help shape the future of travel.
Introducing Reliability & Operations Engineering Trainline is a fast-growing tech company powering world-class digital journeys for millions of customers. Our platform runs primarily on AWS, built on cloud-native architecture, modern CI/CD pipelines, and strong Dev Ops and SRE practices.
The Reliability & Operations Engineering team (Reliability Ops) brings together SRE, Incident Management, and Database Reliability to keep our platform observable, reliable, scalable, and resilient. We partner closely with product engineering teams to enable safe delivery, respond to incidents, and continuously strengthen system reliability.
We’re looking for a mid-level Site Reliability Engineer to help drive this forward. You’ll bring solid production experience, a growth mindset, and a willingness to challenge and be challenged — contributing to platform reliability while developing broader technical ownership with support from senior engineers.
As an SRE at Trainline, you'll be working on...Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platform
Participating in production incident response, supporting investigation, mitigation, communication, and coordinated service restoration
Contributing to post-incident reviews and follow-up actions to improve reliability, scalability, and resilience
Taking part in the SRE on-call rotation
Designing, building, and maintaining observability using metrics, logs, events, and traces to support effective detection and diagnosis
Improving monitoring and alerting by aligning signals to business and customer impact, reducing noise and improving mean time to detection (MTTD)
Ensuring relevant operational data is surfaced quickly and clearly during live incidents
Making informed tooling and technology choices using SRE principles, balancing team and business needs
Supporting AWS-hosted infrastructure and shared platform services using infrastructure-as-code and CI/CD tooling
Collaborating with product engineering teams to ensure services are operationally ready and deployed safely
Advising on reliability and resilience practices
Writing and maintaining reliable, well-structured code and scripts to support reliability and observability goals
Prioritising work effectively and collaborating using agile processes to deliver against team and business goals
Our Tech Stack AWSNew RelicELK stack
Grafana
Incident.io Docker, ECSTerraform
Github Actions We'd love to hear from you if you have...
Experience of SRE concepts such as SLI, SLO and error budgets.
Hands-on experience with observability tooling such as New Relic, Elastic (ELK Stack), Influx, Grafana or similar
Experience working with cloud providers (preferably AWS).Experience troubleshooting Linux operating systems.
Experience of scripting in at least one language (preferably Python)
Understanding of load balancing and reverse proxy concepts, upstream config concepts, upstream health checks, worker & data flow concepts.
Application architecture concepts (threading, queuing, readiness checks, health checks, circuit breakers, timeouts, exponential backoff, throttling).Experience building, maintaining and evolving time series data, retention, cardinality, deviation, moving…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×