Senior Observability Engineer
Listed on 2026-08-29
-
Software Development
DevOps, Cloud Engineer - Software
The Senior Observability Engineer is a hands-on role for someone who has already built production observability capabilities—not only consumed observability platforms as an end user. You will help establish how telemetry is instrumented, collected, processed, visualized, alerted on, and governed across teams, leading early onboardings and shaping practical standards, processes, and ways of working.
Working closely with platform, product, and application teams, you will translate requirements into working telemetry solutions across traces, metrics, and logs, with Open Telemetry as a core and mandatory technology. This role follows a trace-centric approach: metric-based deviations help detect issues, traces are the primary investigation path, and logs are used when additional detail is required. This is an opportunity to shape observability standards and engineering practices at enterprise scale, with direct influence on how teams instrument, operate, and troubleshoot critical services.
The Senior Observability Engineer is a hands-on role focused on building and delivering observability and AIOps as a shared engineering capability and service. You will help establish how observability is designed, implemented, and consumed across teams, leading early onboardings and shaping practical standards, processes, and ways of working.
Working closely with platform, product, and application teams, you will translate requirements into actionable observability solutions, implement and operate core capabilities, and continuously improve how insights, alerting, and automation support reliable services. As a core contributor to the Observability & AIOps Center of Excellence, you will combine deep technical execution with strong ownership of service quality, operational readiness, and capability maturity
What You Will Do- design, deploy, and operate observability and AIOps capabilities for cloud-native and hybrid environments, supporting reliable, production-grade services
- lead the onboarding of teams and services, defining and applying standards for telemetry instrumentation, collection, SLIs, SLOs, and alerting in real-world systems using Open Telemetry
- work hands-on with engineering, SRE, and operations teams to gather requirements and translate them into actionable observability and automation solutions
- install, configure, fine-tune, and maintain Open Telemetry Collectors and telemetry pipelines for applications and platforms
- build and maintain dashboards, alerting, and telemetry flows across distributed tracing, metrics, and logs, leveraging Open Telemetry to deliver meaningful insights and reduce operational noise
- run and evolve observability services that support modern incident detection, investigation, and operational decision-making
- design, configure, and maintain integrations with event management and ticketing systems for alert routing, incident context, and operational workflows
- 8+ years of hands-on experience in observability, SRE, platform, or reliability engineering, including responsibility for production observability capabilities
- proven experience taking observability capabilities from requirements to production-ready solutions; experience limited to dashboards, queries, or observability backends as a consumer is not sufficient
- practical experience instrumenting applications and infrastructure with Open Telemetry to generate traces, metrics, and logs
- strong understanding of modern, trace-centric observability practices beyond traditional log-centric monitoring
- experience deploying, configuring, and operating Open Telemetry Collectors and designing telemetry pipelines is a strong advantage; these will be key hands‑on responsibilities in the role
- experience with cloud-native environments, operating Kubernetes-based workloads
- experience with common observability tool chains and backends (e.g.: Elastic, Grafana, Dynatrace, Datadog, App Dynamics, etc.) - Elastic Stack experience considered a plus
- scripting or automation skills to support repeatable deployment, configuration, validation, and ongoing operations using technologies such as Helm, Terraform, and Ansible
If you have any questions,
check out our FAQ page or call Beata Czyzewskaat / / .
For this vacancy we only accept direct applications.
Diversity is important to us. Therefore, we are looking to receiving applications regardless of any personal background.
#J-18808-LjbffrTo Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: