Program Manager
Listed on 2026-10-05
-
IT/Tech
GPU Cluster Health Repair Strategy Team Lead
The GPU Cluster Health Repair Strategy team is accountable for driving repair strategy, operational execution, and cross-functional governance to keep customer GPU cluster availability at or above 97.5% at all times, where a customer cluster is defined as a shape/location combination. The team establishes KPIs, monitors performance, identifies gaps, launches corrective programs, and partners with engineering, SDE/SRE, data center operations, GSL, CHS, CPV, Warminator, TRS, customer-facing teams, and hardware partners such as NVIDIA and AMD.
This role will lead complex, high-impact program improvements at organizational scale. Defines KPIs, shapes automation, drives senior-leadership alignment, and serves as an escalation point for multi-team issues.
ResponsibilitiesStrategic Program Leadership
- Own strategic repair-transformation programs that materially improve GPU cluster availability, repair velocity, partner accountability, and operational efficiency.
- Define end-to-end program strategy for high-risk availability areas, including long-running repair reduction, proactive risk detection, RMA loop improvement, spare placement optimization, healthy-node custody policy, partner hardware quality feedback, or AI-driven repair insights.
- Translate customer availability risk into strategic priorities, measurable goals, and execution plans.
Availability and Repair Governance
- Drive alignment across customer vertical TPMs, horizontal program TPMs, SDE/SRE, engineering, partner teams, and operations.
- Lead executive-level reviews for high-risk customers, high-risk shapes, or high-risk locations.
- Define governance mechanisms to keep every shape/location cluster above 97.5% availability.
- Advise leadership on SLA/SLO performance, repair productivity, operational bottlenecks, and resourcing gaps.
KPI and Data Framework Ownership
- Define the KPI framework for the team, including business metrics, operational metrics, leading indicators, lagging indicators, and escalation thresholds.
- Build the measurement model for shape/location health, cluster-level availability, repair aging, throughput, RMA loop time, spares sufficiency, hardware failure rate, partner blockers, and tooling efficiency.
- Shape advanced reporting and forecasting requirements with SDE/SRE and data teams.
- Drive consistent use of data to identify risk before customer escalation.
Process and Tooling Transformation
- Lead efforts to reduce manual repair execution through automation, AI-assisted triage, workflow instrumentation, and repair agent tooling.
- Partner with SDE/SRE teams in India, Morocco, and Mexico to prioritize tooling investments that accelerate repair diagnosis, status visibility, escalation, and closure.
- Drive adoption of standardized SOPs, playbooks, partner handoffs, and executive escalation paths.
Partner and Ecosystem Influence
- Lead complex engagements with partners such as NVIDIA and AMD where hardware quality, spare availability, RMA turnaround, or repair playbooks affect customer availability.
- Establish data-backed partner accountability mechanisms.
- Drive corrective action plans where partner performance affects cluster health.
Escalation and Risk Management
- Serve as escalation lead for complex availability risks that span multiple organizations.
- Drive root-cause analysis for recurring repair failures, long-tail repair issues, spare shortages, or process breakdowns.
- Make trade-offs between customer impact, business risk, technical constraints, and operational feasibility.
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements. Range and benefit information provided in this posting are specific to the stated locations only US:
Hiring Range in USD from: $90,100 to $209,500 per annum. May be eligible for bonus and equity. Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business. Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).