Associate Director, Observability and Service Reliability
Job in
Toronto, Jefferson County, Ohio, 43964, USA
Listed on 2026-09-07
Listing for:
Kyndryl
Full Time
position Listed on 2026-09-07
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
Who We Are
At Kyndryl, we run and reimagine the mission‑critical technology systems that drive advantage for the world’s leading businesses. We are at the heart of progress; with proven expertise and a continuous flow of AI‑powered insight, enabling smarter decisions, faster innovation, and a lasting competitive edge. For our people—Kyndryls—that means doing purposeful work that powers human progress. Join us and experience a flexible, supportive environment where your well‑being is prioritized and your potential can thrive.
The Role Enterprise Observability Strategy- Own and mature the enterprise observability and service reliability strategy.
- Define standards for monitoring applications, infrastructure, cloud platforms, networks, endpoints, APIs, databases, middleware, and other critical technology services.
- Establish expectations for metrics, logs, traces, events, synthetic monitoring, real user monitoring, digital experience, service health, and business transaction visibility.
- Create a consistent enterprise approach while allowing teams to use monitoring technologies suited to their platforms and services.
- Identify monitoring gaps, redundant capabilities, excessive alerting, and opportunities to improve visibility.
- Move the organization from traditional monitoring toward proactive, predictive, and automated operations.
- Establish and mature the organization’s service reliability framework.
- Partner with technical service owners to define monitoring requirements for critical applications and services.
- Ensure monitoring reflects the complete service, including application performance, infrastructure, dependencies, integrations, user experience, business transactions, capacity, and failure conditions.
- Define minimum observability requirements based on service criticality and business impact.
- Help teams establish meaningful Service Level Indicators, Service Level Objectives, availability targets, performance thresholds, and health measures.
- Use reliability data, incidents, problem records, capacity trends, and telemetry to identify systemic weaknesses and prioritize improvements.
- Serve as the enterprise technical authority for observability, monitoring, and service reliability.
- Provide architectural guidance to application, infrastructure, cloud, engineering, Dev Ops, SRE, cybersecurity, and operations teams.
- Influence solution design so services are observable, measurable, supportable, and resilient by design.
- Develop enterprise monitoring patterns, reference architectures, standards, and reusable capabilities.
- Guide technical teams in selecting appropriate monitoring methods and technologies for specific platforms and use cases.
- Evaluate emerging observability, AIOps, automation, analytics, and service reliability capabilities for measurable operational value.
- Provide strategic oversight for the enterprise observability and monitoring tool ecosystem.
- Lead the strategy, architecture, governance, adoption, and optimization of major platforms, including Dynatrace, Nexthink, and related enterprise monitoring technologies.
- Ensure monitoring tools operate as an integrated ecosystem rather than isolated platforms.
- Establish standards for instrumentation, tagging, alerting, dashboards, integrations, service mapping, ownership, and data quality.
- Partner with technical teams to maximize platform value while reducing tooling duplication and complexity.
- Manage strategic technology and vendor relationships to ensure observability investments deliver measurable operational value.
- Provide strategic leadership for enterprise use of Dynatrace across applications, infrastructure, cloud, and digital services.
- Drive adoption of application performance monitoring, distributed tracing, real user monitoring, synthetic monitoring, infrastructure monitoring, logs, topology, service health, and intelligent problem detection.
- Partner with technical owners to improve application instrumentation and ensure Dynatrace provides meaningful service visibility.
- Use Dynatrace capabilities to improve root cause identification, dependency…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×