Observability Engineer Specialist
Listed on 2026-09-12
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Cybersecurity, Systems Engineer
Discovery – Group Information Systems - Digital Channels Observability Engineer About Discovery
Discovery's core purpose is to make people healthier and to enhance and protect their lives. We seek out and invest in exceptional individuals who understand and support our core purpose, and whose own values align with those of Discovery. Our fast-paced and dynamic environment enables smart, self-driven people to be their best. As global thought leaders, Discovery is passionate about innovating in order to not only achieve financial success, but to ignite positive and meaningful change within our society.
AboutDigital Channels
Working in a high performance organization that prides itself in attracting the finest talent, we challenge ourselves to find solutions that make a difference in the world. Our environment is always buzzing with energy and smart, motivated people working on finding the best way to move forward.
The Digital Channels team works on dynamic new projects and product enhancements within the web and mobile platforms in order to improve business inefficiencies, gain competitive advantage on our products and ultimately to provide better service to our clients. Using knowledge of the organisation's technology infrastructure and specific software applications, Application Platform Services helps the business to address changes through technologies.
KeyPurpose
The Observability Specialist at Discovery Limited plays a critical role in implementing and maintaining observability solutions that provide end-to-end visibility into system performance and reliability. The role involves deploying and optimising observability platforms, designing telemetry pipelines, and ensuring integration with cloud and container environments. It requires strong technical expertise, collaboration across engineering teams, and the ability to drive continuous improvement in monitoring, alerting, and automation practices.
This position supports incident response, root cause analysis, and the definition of reliability metrics, contributing to a culture of operational excellence and resilience.
- Deploy, configure, and maintain observability platforms such as Prometheus, Grafana, Dynatrace, ELK Stack, and Open Telemetry to ensure reliable and scalable monitoring solutions. Decisions in platform setup directly impact system visibility and operational performance.
- Design and optimize observability systems to provide comprehensive visibility across infrastructure, applications, and services. These design choices influence monitoring accuracy and system resilience.
- Monitor system health and analyze telemetry data, including metrics, logs, and traces, to detect anomalies and performance issues. Insights from this analysis guide proactive remediation and reliability improvements.
- Define and track SLIs, SLOs, and error budgets to align observability practices with reliability objectives. These metrics inform prioritization of engineering tasks and risk management strategies.
- Participate in incident response activities, contributing to root cause analysis and post-incident reviews to improve system resilience. Decisions during incidents affect recovery time and future prevention measures.
- Implement observability solutions in cloud and containerized environments, ensuring coverage across distributed systems. These implementations impact scalability and monitoring consistency.
- Automate observability workflows and integrate observability into CI/CD pipelines using tools such as Jenkins, Git Hub Actions, and Git Lab CI. Automation decisions reduce operational overhead and improve deployment efficiency.
- Write scripts in Python, Bash, or Power Shell to automate observability tasks and support operational workflows. These scripts enhance efficiency and reduce manual intervention.
- Maintain observability documentation, including architecture diagrams, configuration standards, and operational playbooks, to ensure consistency and knowledge sharing across teams.
- Collaborate with cross-functional teams to enhance monitoring coverage, improve alerting accuracy, and promote observability best practices. Engagement includes working with senior engineers, technical leads, and operations managers to align observability strategies with business objectives.
- Internal Stakeholders:
Development teams (junior to senior engineers), Infrastructure and operations teams (technical leads and managers), Site…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).