Vice President, Reliability
Listed on 2026-08-22
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Project Manager
Vice President, Reliability Description -
HP is a technology company founded on the belief that companies should do more than simply generate profit. Having helped shape Silicon Valley decades ago, HP continues to invest deeply in innovation across its broad product portfolio to advance how people live and work. We envision a world where technology improves lives everywhere. At HP, great ideas can come from anyone, anywhere, at any time—and one idea can change the world.
Today, HP is among the world’s most sustainable, just, and inclusive technology companies.
The CTO, Software organization at HP is a global team of thousands of engineers and leaders shaping the future of how software and platform technologies power HP’s product portfolio. The organization is responsible for the vision, strategy, architecture, engineering, and operations of the software and platform technology stacks behind HP’s products. These include modern, AI-native, end-to-end platforms spanning developer experience, data and AI, security, infrastructure, reliability and device software, as well as software stacks for customer-facing HP applications.
Our mission is to deliver magical, AI-enabled experiences to tens of millions of customers worldwide. Beyond powering today’s products, the organization also helps define and shape HP’s next era of growth through technology innovation and active engagement across the software and AI ecosystems.
The Vice President, Reliability is a senior leadership role within the CTO, Software organization responsible for defining and leading the reliability strategy and execution for the shared platform services that power our products and digital experiences. This leader will ensure these platforms are highly available, resilient, observable, performant, and operationally excellent at a global scale.
This role is central to CTO, Software’s platform operating model and to HP’s broader software-led transformation. The Vice President, Reliability will help accelerate engineering velocity, improve customer experience and trust, reduce operational risk, and expand the impact of shared platform capabilities across the company. This leader will establish the standards, mechanisms, and culture required to build and operate reliable systems at scale across a complex, multi-domain engineering environment.
Reporting to the VP, Platform Engineering under the Chief Technology Officer, Software, this executive will serve as a trusted partner to engineering leaders and business stakeholders across HP. This role offers the opportunity to shape how thousands of engineers build, operate, and scale software in one of the world’s most established technology companies—a long-standing Fortune 100 company with more than 28,000 patents, over $55 billion in annual revenue, and 58,000 employees worldwide.
Responsibilities- Define and lead the reliability vision, strategy, operating model, and execution roadmap for shared platform services across developer experience,dataand AI,security,and infrastructure.
- Build and lead a high-performing, inclusive organization of reliability, observability, performance, and resilience leaders and engineers.
- Establish and scale modern site reliability engineering (SRE) practices, including AI-enabled SRE, service-level objectives and indicators, error budgets, production readiness reviews, service maturity models, reliability consulting, and embedded SRE engagement models.
- Lead the observability platform and strategy, including metrics, logs, traces, alerting, dashboards, telemetry standards, service health visibility, and developer-facing operational tooling.
- Own incident management and resilience operations, including major incident command, escalation models, on-call standards, blameless postmortems, disaster recovery, resilience exercises, and fault-injection testing.
- Lead performance and scalability engineering, including load testing, performance profiling, latency optimization, capacity forecasting, and scale validation for critical platform services.
- Drive operational intelligence and automation, including operational metrics, service insights, anomaly detection, and the application of…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).