×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in London, Greater London, W1B, England, UK
Listing for: Trainline
Full Time position
Listed on 2026-09-09
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
About us At Trainline, our purpose is to empower greener travel choices, connecting people and places. Trainline enables millions of travellers to find and book the best value tickets across carriers, fares, and journey options through our highly rated mobile app, website, and B2B partner channels. Great journeys start with Trainline  We’re Europe’s leading independent rail platform, helping millions of travellers find and book the best-value rail and coach journeys across our app, website and partner channels.

Our job is to make the green travel choice the best choice. By building a better train travel experience, we help more people choose rail - creating a positive impact for customers, our business and the planet.

We’re a team of more than 1,000 Train liners from over 50 nationalities, working across London, Paris, Barcelona, Milan, Edinburgh and Madrid. Now is a brilliant time to join us and help shape the future of travel.

Introducing Reliability & Operations Engineering Trainline is a fast-growing tech company powering world-class digital journeys for millions of customers. Our platform runs primarily on AWS, built on cloud-native architecture, modern CI/CD pipelines, and strong Dev Ops and SRE practices.

The Reliability & Operations Engineering team (Reliability Ops) brings together SRE, Incident Management, and Database Reliability to keep our platform observable, reliable, scalable, and resilient. We partner closely with product engineering teams to enable safe delivery, respond to incidents, and continuously strengthen system reliability.

We’re looking for a mid-level Site Reliability Engineer to help drive this forward. You’ll bring solid production experience, a growth mindset, and a willingness to challenge and be challenged — contributing to platform reliability while developing broader technical ownership with support from senior engineers.

As an SRE at Trainline, you'll be working on...Developing an understanding of system architecture, dependencies, and failure modes across the Trainline platform

Participating in production incident response, supporting investigation, mitigation, communication, and coordinated service restoration

Contributing to post-incident reviews and follow-up actions to improve reliability, scalability, and resilience

Taking part in the SRE on-call rotation

Designing, building, and maintaining observability using metrics, logs, events, and traces to support effective detection and diagnosis

Improving monitoring and alerting by aligning signals to business and customer impact, reducing noise and improving mean time to detection (MTTD)
Ensuring relevant operational data is surfaced quickly and clearly during live incidents

Making informed tooling and technology choices using SRE principles, balancing team and business needs

Supporting AWS-hosted infrastructure and shared platform services using infrastructure-as-code and CI/CD tooling

Collaborating with product engineering teams to ensure services are operationally ready and deployed safely

Advising on reliability and resilience practices

Writing and maintaining reliable, well-structured code and scripts to support reliability and observability goals

Prioritising work effectively and collaborating using agile processes to deliver against team and business goals

Our Tech Stack AWSNew RelicELK stack

Grafana

Incident.io Docker, ECSTerraform

Github Actions We'd love to hear from you if you have...

Experience of SRE concepts such as SLI, SLO and error budgets.

Hands-on experience with observability tooling such as New Relic, Elastic (ELK Stack), Influx, Grafana or similar

Experience working with cloud providers (preferably AWS).Experience troubleshooting Linux operating systems.

Experience of scripting in at least one language (preferably Python)
Understanding of load balancing and reverse proxy concepts, upstream config concepts, upstream health checks, worker & data flow concepts.

Application architecture concepts (threading, queuing, readiness checks, health checks, circuit breakers, timeouts, exponential backoff, throttling).Experience building, maintaining and evolving time series data, retention, cardinality, deviation, moving…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary