Vice President, Reliability
Listed on 2026-08-28
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Vice President, Reliability Description - HP is a technology company founded on the belief that companies should do more than simply generate profit. Having helped shape Silicon Valley decades ago, HP continues to invest deeply in innovation across its broad product portfolio to advance how people live and work. We envision a world where technology improves lives everywhere. At HP, great ideas can come from anyone, anywhere, at any time—and one idea can change the world.
Today, HP is among the world’s most sustainable, just, and inclusive technology companies. The CTO, Software organization at HP is a global team of thousands of engineers and leaders shaping the future of how software and platform technologies power HP's product portfolio. The organization is responsible for the vision, strategy, architecture, engineering, and operations of the software and platform technology stacks behind HP’s products.
These include modern, AI-native, end-to-end platforms spanning developer experience, data and AI, security, infrastructure, reliability and device software, as well as software stacks for customer-facing HP applications. Our mission is to deliver magical, AI-enabled experiences to tens of millions of customers worldwide. Beyond powering today’s products, the organization also helps define and shape HP’s next era of growth through technology innovation and active engagement across the software and AI ecosystems.
- Define and lead the reliability vision, strategy, operating model, and execution roadmap for shared platform services across developer experience, data and AI, security, and infrastructure.
- Build and lead a high-performing, inclusive organization of reliability, observability, performance, and resilience leaders and engineers.
- Establish and scale modern site reliability engineering (SRE) practices, including AI-enabled SRE, service-level objectives and indicators, error budgets, production readiness reviews, service maturity models, reliability consulting, and embedded SRE engagement models.
- Lead the observability platform and strategy, including metrics, logs, traces, alerting, dashboards, telemetry standards, service health visibility, and developer-facing operational tooling.
- Own incident management and resilience operations, including major incident command, escalation models, on-call standards, blameless postmortems, disaster recovery, resilience exercises, and fault-injection testing.
- Lead performance and scalability engineering, including load testing, performance profiling, latency optimization, capacity forecasting, and scale validation for critical platform services.
- Drive operational intelligence and automation, including operational metrics, service insights, anomaly detection, and the application of AI to improve reliability engineering and operational response.
- Partner closely with platform, product, security, and engineering executives to embed reliability into architecture, delivery, and operations across the software lifecycle.
- Champion a culture of accountability, engineering excellence, continuous improvement, and customer-centric decision-making.
- Bachelor’s or master’s degree in computer science, engineering, or a related field; a Ph.D. is preferred.
- Current or prior experience operating at the VP level in engineering within an established technology company.
- 15+ years of experience in software, infrastructure, platform, or reliability engineering, including 8+ years leading engineering organizations through senior leadership roles.
- Proven success building and scaling high-performing engineering teams that deliver shared capabilities used across large, complex product or platform environments.
- Experience leading reliability or platform engineering in a global technology company operating at significant scale.
- Experience applying AI and automation to software operations, incident response, and engineering productivity.
- Experience supporting platforms or services that underpin products used by millions of customers and/or large-scale commercial businesses generating more than $100 million in annual revenue.
- Deep expertise in modern cloud-native…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).