×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer

Job in Zürich, 8081, Zurich, Kanton Zürich, Switzerland
Listing for: Open Systems AG
Full Time position
Listed on 2026-08-31
Job specializations:
  • Software Development
Salary/Wage Range or Industry Benchmark: 140000 - 190000 CHF Yearly CHF 140000.00 190000.00 YEAR
Job Description & How to Apply Below
Position: Senior Site Reliability Engineer (80–100%)
Location: Zürich

Senior Site Reliability Engineer (80–100%) Your Mission

We are seeking a highly motivated and talented Senior Site Reliability Engineer to join our growing team. As an SRE, you will be a key driver in building automation, reducing operational toil, and continuously improving how our services are operated, while ensuring reliability, scalability, and accurate SLA measurement.

You will combine strong software engineering and system architecture skills with an operations mindset to improve how our services are built and run h of your work centers on the Mission Control (MC) platform— the operational backbone our engineers rely on to run customer services in production, and our point of interaction with customers, running 24/7 around the globe. To improve Mission Control, we build services such as an automation framework, a self‑service platform, and a service‑level monitoring system that eliminate toil and repetitive tasks, enable customers, and make operations more reliable and efficient.

Your responsibilities will include:

  • Building Operational Automation:
    Design, build, and evolve the automation framework and tooling that powers the MC platform — primarily in Golang— with a strong focus on maintainability, scalability, and reliability. Beyond the framework itself, you will build automation workflows on top of the platform that reduce operational toil and improve day‑to‑day efficiency for the engineers and operators who run our services.
  • Developing Self‑Service APIs:
    Build and maintain the service APIs and self‑service operational tooling of the MC platform that enable customers and teams to safely and efficiently operate services in production without manual intervention.
  • Applying Site Reliability Engineering (SRE) Principles:
    Define, implement, and continuously improve Service Level Indicators (SLIs), Service Level Objectives (SLOs), error budgets, and SLA measurements so that reliability is measurable and actionable across the MC platform and the services it operates.
  • Owning Reliability and Operations Initiatives:
    Take ownership of reliability, automation, and Mission Control projects, driving them independently from problem identification through implementation and long‑term operation.
  • Collaborating with AI Engineering:
    Work closely with our AI team and tooling, integrating AI‑assisted capabilities into our automation and operational workflows.
  • Incident Response and Learning:
    Participate in incident response and the on‑call rotation, leading root cause analysis and driving sustainable corrective and preventive actions.

This position will give you the opportunity to lead midsize to large automation and reliability initiatives, from early concept and design through production deployment and ongoing operations. You will collaborate closely with software engineers, platform teams, and product owners to identify and implement improvements across our services and operational practices. As part of the role, you will spend at least one day per week working within Mission Control, staying close to the day‑to‑day reality of how our services are run.

Your

Qualifications

You are a senior engineer with a strong foundation in software development and system architecture, combined with a clear interest in operations and reliability. You are comfortable taking ownership, learning quickly, and driving change without needing detailed direction.

Required skills and qualifications:

  • Strong software engineering background, ideally with production‑grade Go (Golang) experience — proficiency in another modern language is fine if you are a quick learner and willing to adapt, as Go is our primary language.
  • Solid understanding of distributed systems and scalable architecture.
  • Proven experience designing, building, and operating services and their APIs (e.g. REST, gRPC) in production.
  • Experience operating production systems, including incident response, on‑call, and root cause analysis.
  • Experience with SRE, Dev Ops, platform engineering, or reliability‑focused roles.
  • Hands‑on experience with infrastructure and operations tooling, such as:
    • Kubernetes
    • Terraform / Infrastructure as Code
    • Git Ops principles and CI/CD tooling
    • Prometheus,…
Position Requirements
10+ Years work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary