×
Hier anmelden um sich kostenlos auf Stellen zu bewerben oder Stellenanzeigen aufzugeben. X

Senior Site Reliability Engineer​/SRE – Kubernetes & Hybrid Cloud; m​/f​/d

in 10115, Berlin, Berlin, Deutschland
Unternehmen: FACT-Finder
Vollzeit position
Verfasst am 2026-09-21
Berufliche Spezialisierung:
  • IT/Informationstechnik
    Site Reliability Ingenieur/in, Cloud Computing: IT-Infrastruktur & Betrieb
Gehalts-/Lohnspanne oder Branchenbenchmark: 110000 - 150000 EUR pro Jahr EUR 110000.00 150000.00 YEAR
Stellenbeschreibung
Stellenbezeichnung: Senior Site Reliability Engineer / SRE – Kubernetes & Hybrid Cloud (m/f/d)

Introduction

At a glance

  • Location &workmodel:

    Berlin, hybrid
  • Tech stack:

    Kubernetes on our own servers, Harvester (Kube Virt), Argo CD/Flux, Prometheus/Grafana, Longhorn/Ceph
  • Team:

    A growing SRE team – you report to our CTPO for now and to the Team Lead SRE we're hiring next; two system administrators in Pforzheim run the physical hardware
  • Process:

    Intro call
    · take-home task (~2h)
    · 90-min tech interview with our developers
    · leadership conversation
    · meet the team
  • Languages:

    Fluent English required;
    German is a plus, not a must

Why this role is special

Most SRE jobs today mean clicking around a managed cloud console. This one doesn't. We run our own hardware in Frankfurt and are building a modern private cloud platform on Kubernetes and Harvester – on-prem by default, with elastic burst into the public cloud and the option to go cloud-only later. You won't inherit a finished SRE practice: you'll help define it, side by side with our Berlin development teams – and you won't do it alone, a Team Lead SRE hire is coming next.
SRE here is an enabling discipline: you build what our developers need to ship reliably, while two system administrators in Pforzheim run the physical hardware. And the impact is direct – our product discovery technology powers more than 2,000 European online shops (Intersport, SPAR, Douglas and more), handling billions of shopper queries a year. When product discovery is slow or down, our customers lose revenue in real time.

Your

first 90 days

You get to know both products, join the on-call rotation with a buddy, and own your first reliability topic – SLOs for one product, alerting that actually helps at 3 a.m., or automating away a piece of toil. By day 90 you've shipped visible improvements and know where you want to take the platform next.

Your mission
  • Define and own SLOs, SLIs and error budgets; drive data-informed reliability decisions
  • Lead incident response end-to-end: fast detection, clear communication, blameless postmortems – and reduce whole classes of incidents structurally, not case by case
  • Eliminate toil through automation and Git Ops; evolve our observability (metrics, logs, traces, alerting, runbooks) across two different stacks
  • Help build our custom Kubernetes operator (CRDs) that makes stateful search clusters declarative, self-healing and safely upgradable – and roll out the auto-scaling (HPA/VPA, KEDA, cluster auto scaler) today's architecture makes hard
  • Plan capacity, performance and cost across on-premises and cloud – including the large-catalogue and peak-season loads our merchants care about – and use AI tools wherever they measurably speed up diagnosis and operations
Your profile

Must-haves:

  • Kubernetes in production – built, not just used: you've set up and maintained clusters on your own servers (e.g.kubeadm, RKE2, k3s) and know cluster lifecycle and upgrades – managed-only experienceisn'tenough for this role
  • Lived SRE practice: SLOs, error budgets, incident management,on-call
  • Hands‑on experience with

    GitOpsor comparable infrastructure/deployment automation– experience with Argo CD or Flux is a strong plus
  • Solid observability skills– metrics, logs, traces, alerting that people trust
  • A strong automation instinct– you'd rather fix a problem's cause than repeat its workaround
  • A collaborative, enabling mindset– you see SRE as a service to our developers: you ask what they need, discuss trade-offs openly, and don't fall in love with your own solution

Nice-to-haves (genuinely optional –we'll teach you the rest):

  • Harvester,Kube Virt, vSphere/ESXi, Open Stack or similar virtualization/HCI platforms
  • Container storage (Longhorn, Ceph) and datacenter networking (load balancing, ingress, VLAN)
  • Auto-scaling (HPA, VPA, KEDA, cluster auto scaler) and capacity/cost planning
  • Experience…
Stellen-Anforderungen
10+ Jahre Berufserfahrung
Um Jobs auf dieser Seite anzusehen und sich zu bewerben, die Bewerbungen aus Ihrem Standort oder Land akzeptieren, klicken Sie unten auf den Button, um eine Suche zu starten.
(Wenn dieser Job tatsächlich in Ihrem Zuständigkeitsbereich liegt, verwenden Sie möglicherweise einen Proxy oder VPN, um auf diese Seite zuzugreifen. Um weiterzukommen, sollten Sie Ihre Verbindung zu einem anderen Mobilgerät oder PC wechseln).
 
 
 
Suchen Sie hier nach weiteren Stellen:
(nach Beruf, Fähigkeit)
Standort
Suchradius erweitern (Meilen)
0
200
Filter
Mindest-Bildungsgrad für die Stelle
Mindest-Berufserfahrung für die Stelle
Veröffentlicht in den letzten:
Gehalt