×
Register Here to Apply for Jobs or Post Jobs. X

Ingénieur Senior Plateforme d'Observabilité

Job in 3000, Bern, Canton de Berne, Switzerland
Listing for: Jobup
Full Time position
Listed on 2026-07-09
Job specializations:
  • Software Development
    DevOps, Cloud Engineer - Software
Salary/Wage Range or Industry Benchmark: 100000 - 130000 CHF Yearly CHF 100000.00 130000.00 YEAR
Job Description & How to Apply Below
Position: Ingénieur Senior Plateforme d'Observabilité (80-100%)

Senior Observability Platform Engineer (80-100%)

Location:

Zurich / Bern

We are seeking a highly skilled and experienced Senior Platform Observability Engineer to join our team. In this role, you will be responsible for ensuring the reliability, scalability, and efficiency of our core observability infrastructure that supports our engineering teams and customer‑facing portal. Your work will include evolving these systems and participating in fostering adoption of observability best‑practices in the organization.

Key Responsibilities
  • Configure, operate, and enhance our observability platforms and frameworks (Clickhouse, Thanos, Loki, Tempo, Open Telemetry Collector + custom processors).
  • Continuously improve and drive organization‑wide adoption of observability best‑practices, ensuring comprehensive monitoring, logging, and tracing.
  • Develop and maintain automated solutions for monitoring, alerting, and incident response.
System Optimization
  • Collaborate with engineering teams to understand their needs and provide robust, scalable solutions utilizing the observability platform.
  • Optimize system performance and ensure high availability through proactive monitoring and maintenance.
  • Develop and implement strategies for cost optimization, capacity planning, and performance tuning.
Innovation and Improvement
  • Stay up‑to‑date with the latest industry trends, tools, and technologies to drive continuous improvement.
  • Experiment with and implement new tools, especially around observability and telemetry, to enhance platform capabilities.
  • Evaluate and integrate Open Telemetry Collector where beneficial to enhance telemetry data collection and analysis.
Required Skills and Experience
  • Observability Platforms:
    Proven track record in managing at least one of the following observability stacks:
    Thanos, Mimir, Cortex, Tempo, Loki, or Clickhouse; with the ability to configure, operate, and improve these systems.
  • Kubernetes:
    Deep understanding of Kubernetes architecture and hands‑on experience in managing resources on clusters.
  • Helm:
    Experience in writing and maintaining Helm charts, and understanding third‑party charts to deploy and manage Kubernetes resources efficiently.
  • Git Ops:
    Experience in continuous delivery and Git Ops practices (version control, CI/CD pipelines).
  • Agentic Development:
    Hands‑on experience using agentic AI workflows (e.g., Git Hub Copilot, Claude Code, Cursor, or similar) to accelerate day‑to‑day engineering.
  • Docker:
    Expertise in containerization, orchestration, and optimization of Docker workloads.
Desirable Skills
  • Coding

    Experience:

    Coding knowledge in Golang or a similar language.
  • Open Source:
    Contributor to open source project written in Golang or a similar language.
  • Open Telemetry Collector:
    Knowledge of the Open Telemetry Collector or direction contribution to project.
  • AI for Observability:
    Interest in applying AI/ML to the observability domain like anomaly detection on metrics and logs, automated root‑cause analysis, alert noise reduction and correlation, and natural‑language querying over telemetry.
Soft Skills
  • Quick Learner:
    Ability to quickly grasp new concepts and technologies, adapting to the evolving needs of the organization.
  • Communication:
    Excellent communication skills, with the ability to convey complex technical concepts to both technical and non‑technical stakeholders.
  • Customer Focus:
    Keen awareness of customer needs and the impact of platform operations on both internal engineering teams and external users.
  • Collaborative Mindset:
    Strong ability to work collaboratively in cross‑functional teams, contributing to a culture of continuous improvement and innovation.
Education and Experience
  • Bachelor’s degree in Computer Science, Information Technology, or related field (or equivalent experience).
  • 5+ years of experience in platform engineering, site reliability engineering, or a related role.
  • Demonstrated experience in managing large‑scale infrastructures and observability platforms (such as Thanos, Mimir, Cortex, Tempo, Loki, Clickhouse).
  • Technical Expertise.
  • Observability Platform Operations.
  • You are excited by the prospect of managing more than 20 TB of telemetry data per day, originating from a fleet of 10 000+ nodes (including linux hosts, k8s clusters, VMs).
Equal Opportunity

Come as you are! We search for amazing people of diverse backgrounds, experiences, abilities, and perspectives. Open Systems welcomes and encourages diversity in the workplace regardless of race, gender, religion, age, sexual orientation, disability, or veteran status.

#J-18808-Ljbffr
Position Requirements
10+ Years work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary