Kubernetes Platform / DevOps Engineer
Job in
Zürich, 8081, Zurich, Kanton Zürich, Switzerland
Listed on 2026-08-16
Listing for:
Michael Bailey Associates
Full Time
position Listed on 2026-08-16
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Job Description & How to Apply Below
We are looking for an experienced Kubernetes Platform / Dev Ops Engineer to operate, maintain and continuously improve a production Kubernetes-based platform and the managed services running on it.
You will play a key role in ensuring the platform remains highly available, scalable, secure and resilient
, while driving automation, observability and operational improvements across the environment.
This is a hands‑on operational role suited to someone with strong Kubernetes, Linux, automation and production infrastructure experience.
Key Responsibilities- Operate, monitor and continuously improve a Kubernetes-based platform and managed services including MongoDB, PostgreSQL and Kafka
. - Ensure high availability, scalability and resilience through proactive monitoring, alerting and incident management
. - Perform and automate Day 2 operations
, including upgrades, patching, backup and restore, capacity management and maintenance. - Investigate production incidents and perform detailed error analysis and Root Cause Analysis (RCA).
- Implement and maintain observability solutions using Prometheus, Grafana and ELK/Open Search
. - Develop and maintain automation, runbooks and operational processes to improve platform reliability and efficiency.
- Work closely with engineering teams to continuously improve the platform and operational tooling.
- Support production environments and participate in an on-call rotation
, ensuring reliable 24/7 operation.
- Degree in Computer Science or equivalent education, combined with several years of professional experience in IT operations.
- Strong hands‑on experience with Kubernetes and operating container-based platforms and services.
- Solid expertise in observability
, particularly Prometheus, Grafana and ELK/Open Search. - Strong experience with Linux/system engineering in production environments.
- Experience with automation and Infrastructure as Code
, using technologies such as: - Ansible
- Terraform
- Experience with CI/CD pipelines
, Git Lab, Git and Docker. - Operational experience with at least one of MongoDB, PostgreSQL or Kafka
. - Strong troubleshooting and structured problem‑solving skills, including incident investigation and RCA.
- Ability to quickly understand existing infrastructure and make improvements from day one.
- Comfortable working independently while collaborating effectively within an agile engineering environment.
- Resilient and calm under pressure, particularly when dealing with production incidents.
- Kubernetes certifications such as CKA or CKAD
, or equivalent practical experience. - Experience with CNCF technologies such as ArgoCD, Velero, Cilium or Kyverno
. - Knowledge of security hardening, secrets management and RBAC/IAM
. - Experience with IT service management / ITIL and 3rd-level support.
- Previous experience within Service Provider, Telco or Managed Service environments
. - Willingness to participate in on-call / 24x7 support
.
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
Search for further Jobs Here:
×