×
Hier anmelden um sich kostenlos auf Stellen zu bewerben oder Stellenanzeigen aufzugeben. X

Senior Site Reliability Engineer (m​/f​/d

in 67657, Kaiserslautern, Rheinland-Pfalz, Deutschland
Unternehmen: TOPdesk
Vollzeit position
Verfasst am 2026-08-20
Berufliche Spezialisierung:
  • IT/Informationstechnik
    Site Reliability Ingenieur/in, Cloud Computing: IT-Infrastruktur & Betrieb, Systemingenieur
Gehalts-/Lohnspanne oder Branchenbenchmark: 90000 - 150000 EUR pro Jahr EUR 90000.00 150000.00 YEAR
Stellenbeschreibung
Stellenbezeichnung: Senior Site Reliability Engineer (m/f/d)
Location: Kaiserslautern

Company Description

TOPdesk builds service management software used across education, healthcare, government, and manufacturing. We are 700+ colleagues in 8 offices worldwide. Founded over 30 years ago, we serve more than 10 million users worldwide and have been helping organisations deliver better services ever since.

We are an open, collaborative organisation with little hierarchy — people own their work end to end and are trusted to make the decisions that matter. We are reinventing ITSM and ESM for the agentic era, building AI agents our customers can trust, and this role is part of that.

Job Description

About the role

Our Azure SaaS estate keeps service management running for thousands of organisations worldwide, under SLA-backed 24/7 availability. As a Senior Site Reliability Engineer, you own the reliability of that estate as an engineering problem — you set the SLOs, engineer out the toil behind them, and make the platform faster to change and cheaper to operate without trading away resilience.

You sit in the SaaS infrastructure function, working alongside cloud engineering and the product squads shipping to production. You bring our AI-native ways of working into reliability: agents and bounded automation with observability, approvals, containment, and rollback — self-healing systems, not runbooks worked by hand.

What this is not A ticket-driven, break-fix ops role kept away from the code. This is reliability as engineering — you own SLOs and error budgets, automate what you repeat, and design the platform to recover itself rather than reacting incident by incident.

The team We are a group of social technicians who value transparency, open feedback, and a healthy work-life balance — and who treat reliability as a shared, measurable objective, not a firefight.

What you'll own
  • SLOs and error budgets. Define and own service-level objectives across the Azure (and potentially multi-cloud) estate, and use error budgets to steer the balance between shipping change and protecting reliability.
  • Toil elimination and self-healing automation. Identify toil, classify it, and engineer it out — feeding self-healing automation and your findings into the reliability roadmap. Standupanagent-basedsupportlayerthatownsrecurringtoilandcontinuouslyfeedsimprovementsbackintoreliability.
  • Observability consolidation. Standardise metrics, alerting, and tracing across all datacenters, close coverage gaps on cloud workloads, and measurably reduce the alert-to-incident ratio from baseline.
  • Incident response and blameless postmortems. Lead incidents to resolution, run blameless postmortems, and turn every learning into a durable fix or an automation candidate.
  • Reliability of releases. Harden CI/CD and progressive delivery — canaries, safe rollouts, automated rollback — so change velocity and reliability rise together.
  • Capacity and performance. Model capacity, load-test critical paths, and keep the platform within its performance envelope as it scales across regions.
  • AI-native reliability. Bring agents and bounded automation — with observability, approvals, containment, and rollback — into detection, diagnosis, and remediation.
  • Runbooks that get used. Every alert links to a runbook; every runbook links to an automation candidate. You leave things more legible than you found them.
  • Capacity and cost forecasting. Own capacity and cost planning across the multi-cloud estate, model usage and growth trends, and forecast short and long term infrastructure needs so spend and scaling decisions stay ahead of demand rather than reacting to it.
How you approach the work
  • Automate what you repeat — if you have done it manually twice, the third time is a design problem.
  • Measure before optimising: SLOs, baselines, and dashboards before opinions.
  • Design for failure — assume things break, and make recovery automatic and observable.
  • Consultative, not gatekeeping: you pair with product engineering teams and transfer knowledge as you go.
  • Treat cost and reliability as joint objectives, not a forced trade-off.
  • Pro-active collaboration with product teams. You are involved in the early phases of product development, including design to help the teams make…
Stellen-Anforderungen
10+ Jahre Berufserfahrung
Bitte beachten Sie, dass derzeit keine Bewerbungen aus Ihrem Zuständigkeitsbereich für diese Stelle über diese Jobseite akzeptiert werden. Die Präferenzen der Kandidaten liegen im Ermessen des Arbeitgebers oder des Personalvermittlers und werden ausschließlich von diesen bestimmt.
Um nach Stellen zu suchen, sie anzusehen und sich zu bewerben, die Bewerbungen aus Ihrem Standort oder Land akzeptieren, klicken Sie hier, um eine Suche zu starten:
 
 
 
Suchen Sie hier nach weiteren Stellen:
(nach Beruf, Fähigkeit)
Standort
Suchradius erweitern (Meilen)
0
200
Filter
Mindest-Bildungsgrad für die Stelle
Mindest-Berufserfahrung für die Stelle
Veröffentlicht in den letzten:
Gehalt