Vice President of IT Observability & Resiliency
Listed on 2026-07-13
-
IT/Tech
SRE/Site Reliability, AI Business & Operations, Cloud Computing: Infrastructure & Operations
Optum Tech is a global leader in health care innovation. Our teams develop cutting‑edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care's most complex challenges. Your contributions here have the potential to change lives.
Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.
Join the Global Technology Services (GTS) team as our Vice President of IT Observability & Resiliency. In this high‑impact executive role, you will define and drive our enterprise observability strategy across hybrid‑cloud and on‑premise platforms, moving the organization from reactive monitoring to proactive, predictive operations. You will champion the transformation of our enterprise reliability by building a centralized, "always‑on" observability data platform that integrates metrics, events, logs, and traces (MELT).
Utilizing cutting‑edge AIOps, telemetry frameworks (such as Open Telemetry), and automated self‑healing capabilities, you will lead a highly talented team of platform and reliability engineers to deliver end‑to‑end visibility and world‑class system availability for our global healthcare ecosystem.
You'll enjoy the flexibility to work remotely
* from anywhere within the U.S. as you take on some tough challenges.
For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities:- Define and drive the enterprise observability strategy - metrics, events, logs, and traces ('MELT') - across hybrid‑cloud and on‑premise platforms, successfully transitioning the organization from reactive monitoring to proactive, predictive operations
- Own the architecture, standardization, and consolidation of enterprise observability platforms, including APM, logging, tracing, alerting, and dashboards, to establish robust standard telemetry frameworks
- Strategically deploy artificial intelligence and predictive modeling to solve complex infrastructure anomalies, proactively unlock platform performance opportunities, and deliver measurable business value
- Champion predictive technologies as core drivers of operational reliability, actively shaping how automated telemetry, self‑healing systems, and proactive observability are developed and executed
- Partner closely with SRE and operations teams to reduce MTTR, build single‑source‑of‑truth telemetry, and enable "always‑on" enterprise reliability outcomes, including availability, resiliency, and recovery readiness
- Integrate modern, enterprise‑approved AI and machine learning tools into operational workflows, setting continuous learning goals and building organizational competencies to accelerate adoption of automated operations
- Ensure compliance with security, resilience, and regulatory requirements, establishing strict telemetry data governance, including PII/PHI controls and access management
- Champion the ethical, transparent, and accountable use of artificial intelligence and machine learning technologies throughout the observability and telemetry lifecycle
You’ll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.
Required Qualifications:- 15+ years of experience in infrastructure, Site Reliability Engineering (SRE), or platform engineering leadership
- 5+ years of experience directing large‑scale, complex cloud or enterprise infrastructure transformation initiatives
- 5+ years of experience managing, scaling, or architecting enterprise observability, monitoring, or APM platforms
- 3+ years of experience applying AI/ML technologies, such as AIOps, anomaly detection, or predictive analytics, to optimize enterprise IT operations
- Demonstrated experience leading cross‑functional engineering teams and influencing senior executive stakeholders at the CIO/CTO level
- Bachelor's or Master's degree in Computer…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).