HPC Platform Product Lead
Listed on 2026-09-11
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure
Position Summary
Owns the end‑to‑end success of Zoetis’ VMRD High Performance Computing (HPC) environment as both a product and a platform. Sets vision and roadmap, defines service offerings and governance, and provides technical leadership for architecture, SLURM scheduling, reliability, and sustainable growth. Partners with scientific/analytics users and business leaders to translate workload demand into a prioritized backlog and measurable value narrative. Ensures the environment delivers the right performance, cost, and risk posture through capacity planning, operational excellence, and data lifecycle strategy.
Responsibilities- Product ownership & value management (25%):
Own product vision, roadmap, service catalog, and backlog with clear acceptance criteria. Translate utilization, throughput, and outcomes into investment asks, funding recommendations, and executive‑ready value stories (showback/chargeback as applicable). - HPC architecture & technical standards (20%):
Define target‑state architecture and standards across compute, storage, network, security, and tooling; drive multi‑year evolution planning. - SLURM scheduling leadership & L3 support (15%):
Own scheduling strategy (partitions/queues, QoS, fairshare, accounting, reservations); lead configuration, troubleshooting, and job efficiency analysis; provide technical backup as needed. - Operations, reliability, and observability (20%):
Drive operational excellence via monitoring/alerting, dashboards, SLO/SLI practices, incident/problem follow‑up, runbooks, and continuous improvement. - Capacity planning, procurement, and lifecycle management (10%):
Plan and execute compute and storage growth; manage allocations; coordinate procurement and lifecycle refresh. - Governance, access, and user enablement (10%):
Establish onboarding, entitlements, RBAC, and auditability; improve user support processes (ticketing, escalations, knowledge base); publish workload best practices (CPU/GPU optimization, execution standards, data movement).
Required:
Bachelor’s degree in computer science, engineering, information systems, or equivalent practical experience. 7+ years of experience in infrastructure engineering/architecture, platform engineering, or adjacent technical product roles. 5+ years of experience in HPC environments (on‑prem, cloud, and/or hybrid). Demonstrated experience partnering with scientists/engineers to translate workload needs into prioritized platform capabilities and measurable value stories.
Preferred:
Master’s degree in a relevant field; experience operating HPC in regulated or highly governed environments; product owner/Agile delivery experience (e.g., backlog management, acceptance criteria, road‑mapping).
- Strong SLURM administration skills: partitions, queues, QoS, fairshare, accounting, reservations; job monitoring and troubleshooting; job efficiency analysis; user guidance.
- HPC infrastructure knowledge:
Linux fundamentals, compute (CPU/GPU), high‑speed networking, shared/parallel storage concepts, capacity/performance planning. - Automation/scripting:
Bash/Python; infrastructure‑as‑code/automation tooling where applicable. - Observability: utilization telemetry, job analytics, dashboards/alerting for node health and storage utilization; actionable operational reporting.
- Data lifecycle and protection: retention policies, tiered storage, archive/backup strategy, restore testing, and operational runbooks.
- Product and delivery skills: requirements discovery, value storytelling, roadmap/backlog prioritization, stakeholder management, and communication of technical trade‑offs and ROI.
Position Requirements
Primarily an office/remote knowledge‑work role; prolonged periods of sitting and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).