Senior Vice President, Product Reliability Engineering
Listed on 2026-07-13
-
IT/Tech
SRE/Site Reliability
Job Description
The Senior Vice President (SVP) of Product Reliability Engineering (PRE) is a senior technology leader responsible for ensuring that Visa's products, platforms, and services are engineered and operated to meet world‑class standards of availability, resiliency, performance, and disaster recovery readiness. This is not a traditional IT operations role. It is an engineering‑first leadership role accountable for reliability strategy, frameworks, tooling enablement, and executive incident leadership across a complex global portfolio.
Key Responsibilities Reliability Engineering Leadership- Define and lead enterprise‑wide reliability engineering strategy, frameworks, and standards (e.g., SLOs, error budgets, reliability governance).
- Drive consistent adoption of reliability practices across product and platform engineering teams.
- Operate effectively in a highly matrixed environment with strong cross functional influence.
- Incorporate GenAI tooling and agentic capabilities to strengthen reliability outcomes across monitoring/alerting, rapid incident response, change management/testing, and Dev Ops/deployment processes.
- Drive adoption of agentic workflows that reduce operational load and accelerate restoration, triage, diagnostics, and remediation execution.
- Advance GenAI capabilities that improve deployment safety and testing rigor, enabling faster change with maintained operational resilience.
- Establish expectations for AI Observability and instrumentation for AI systems, partnering with relevant engineering teams to operationalize monitoring and lifecycle management.
- Partner across technology teams to accelerate productivity through GenAI, including enabling defined classes of work to be implemented end to end by AI agents where appropriate.
- Provide executive leadership during high severity incidents, including rapid decisioning, cross‑team mobilization, and stakeholder communication.
- Ensure strong post‑incident learning loops and systemic engineering remediation.
- Own reliability enabling platforms and tooling (automation, resilience validation, deployment safety mechanisms, reliability guardrails) used by engineering teams to operate safely at scale.
- Lead performance and scalability engineering initiatives across products and platforms to ensure systems meet latency, throughput, and efficiency requirements.
- Partner with product, platform, and infrastructure leaders to embed reliability as a design requirement (not a downstream function).
- Establish operating rhythms and governance mechanisms that balance speed of delivery with reliability and risk management.
- Own ITDR strategy and operational readiness to ensure systems meet recovery objectives and resilience requirements.
- Lead resilience planning and execution across critical services and dependencies.
Basic Qualifications
- 15 or more years of relevant work experience with a Bachelor's Degree or at least 12 years of work experience with an Advanced degree (e.g. Masters/MBA/JD/MD) or a minimum of 10 years of work experience with a PhD.
Preferred Qualifications
- 15+ years of experience in large scale, mission critical environments.
- Experience leading global engineering or infrastructure teams.
- Expertise in ITDR, resilience architecture, and reliability frameworks.
- Experience in regulated / high availability industries.
- Deep knowledge of distributed systems and resilience patterns.
- Experience incorporating GenAI tooling and agentic capabilities into operational and reliability workflows (especially monitoring/alerting, response, change management/testing, and Dev Ops/deployment).
- Familiarity with AI Observability and instrumentation approaches for AI systems.
The estimated salary range for this position is $$292,200 to $$575,000 USD per year, which may include potential sales incentive payments (if applicable). Salary may vary depending on job‑related factors which may…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).