×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Greater London, London, Greater London, W1B, England, UK
Listing for: Incite Insight
Full Time position
Listed on 2026-08-29
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 70000 - 90000 GBP Yearly GBP 70000.00 90000.00 YEAR
Job Description & How to Apply Below
Location: Greater London

Site Reliability Engineer Incite Insight London, England, United Kingdom

Site Reliability / Platform Engineer We are recruiting for a growing technology infrastructure business that is building a new operational capability to support large-scale, high-performance computing environments.

This is an excellent opportunity for an experienced Site Reliability Engineer or Platform Engineer who enjoys automating things rather than repeatedly fixing them manually.

The role sits at the intersection of infrastructure, operations and software engineering. You will use Python and modern automation techniques to improve reliability, reduce manual workload and make incident response faster and more effective.

What you'll be doing You'll build Python-based automation around incident management, operational runbooks and routine infrastructure tasks.

You'll integrate monitoring, infrastructure and ITSM platforms through APIs, helping improve the quality of alerts through better correlation, enrichment, suppression and deduplication.

You'll also develop internal tools, command-line utilities, dashboards and potentially Chat Ops capabilities that allow operational teams to resolve issues more quickly.

A major part of the role will be taking existing operational processes and asking:
Why are we still doing this manually? You'll then design a safe, controlled and auditable way of automating it.

What we're looking for You should have good commercial experience in Site Reliability Engineering, Platform Engineering, Dev Ops or production infrastructure operations, together with strong hands-on Python automation skills.

You'll also need experience with:
Monitoring and observability tools such as Prometheus, Grafana or similar Production incident management and/or on-call environments Automating operational runbooks and repetitive infrastructure processes APIs and systems integration Version-controlled automation and operational tooling

Experience with any of the following would be particularly useful:

Service Now, Halo, Jira Service Management, Open Telemetry, distributed tracing, Slack/Teams automation, data centre or colocation environments, GPU infrastructure, DCIM, IPAM, virtualisation platforms or LLM-assisted operational automation.

This is not an AI/ML development position. We're looking for someone who understands production infrastructure and can use software engineering and automation to make that infrastructure more reliable.

You'll be joining a growing organisation where you'll have considerable autonomy and the opportunity to help shape the SRE and operational automation capability rather than simply inherit an established environment.

Salary: TBCLocation / hybrid working: TBC

#J-18808-Ljbffr
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary