HPC Platform Product Lead
Listed on 2026-07-25
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Support, SRE/Site Reliability
Owns the end-to-end success of Zoetis' VMRD High Performance Computing (HPC) environment as both a product and a platform: sets vision and roadmap, defines service offerings and governance, and provides technical leadership for architecture, SLURM scheduling, reliability, and sustainable growth. Partners with scientific/analytics users and business leaders to translate workload demand into a prioritized backlog and measurable value narrative. Ensures the environment delivers the right performance, cost, and risk posture through capacity planning, operational excellence, and data lifecycle strategy.
POSITIONRESPONSIBILITIES POSITION SUMMARY
Owns the end-to-end success of Zoetis' VMRD High Performance Computing (HPC) environment as both a product and a platform: sets vision and roadmap, defines service offerings and governance, and provides technical leadership for architecture, SLURM scheduling, reliability, and sustainable growth. Partners with scientific/analytics users and business leaders to translate workload demand into a prioritized backlog and measurable value narrative. Ensures the environment delivers the right performance, cost, and risk posture through capacity planning, operational excellence, and data lifecycle strategy.
Productownership & value management (25%)
- Own product vision, roadmap, service catalog, and backlog with clear acceptance criteria.
- Translate utilization/throughput/outcomes into investment asks, funding recommendations, and executive-ready value stories (showback/chargeback as applicable).
- Define target-state architecture and standards across compute, storage, network, security, and tooling; drive multi-year evolution planning.
- Own scheduling strategy (partitions/queues, QoS, fairshare, accounting, reservations); lead configuration, troubleshooting, and job efficiency analysis; provide technical backup as needed.
- Drive operational excellence via monitoring/alerting, dashboards, SLO/SLI practices, incident/problem follow-up, runbooks, and continuous improvement.
- Plan and execute compute and storage growth; manage allocations; coordinate procurement and lifecycle refresh.
- Establish onboarding/entitlements/RBAC and auditability; improve user support processes (ticketing, escalations, knowledge base) and publish workload best practices (CPU/GPU optimization, execution standards, data movement).
Required:
- Bachelor's degree in computer science, Engineering, Information Systems, or equivalent practical experience.
- 7+ years of experience in infrastructure engineering/architecture, platform engineering, or adjacent technical product roles.
- 5+ years of experience in HPC environments (on-prem, cloud, and/or hybrid).
- Demonstrated experience partnering with scientists/engineers to translate workload needs into prioritized platform capabilities and measurable value stories.
- Master's degree in a relevant field.
- Experience operating HPC in regulated or highly governed environments.
- Product Owner / Agile delivery experience (e.g., backlog management, acceptance criteria, road mapping).
- Strong SLURM administration skills: partitions/queues, QoS, fairshare, accounting, reservations; job monitoring and troubleshooting; job efficiency analysis; user guidance.
- HPC infrastructure knowledge:
Linux administration fundamentals, compute (CPU/GPU), high-speed networking, shared/parallel storage concepts, and capacity/performance planning. - Automation/scripting (e.g., Bash/Python) and operational tooling; infrastructure-as-code/automation tooling where applicable.
- Observability: utilization telemetry, job analytics, dashboards/alerting for node health and storage utilization; actionable operational reporting.
- Data lifecycle and protection: retention policies, tiered storage, archive/backup strategy, restore testing, and operational runbooks.
- Product and delivery skills: requirements discovery, value storytelling, roadmap/backlog prioritization, stakeholder management, and communication of technical trade-offs and ROI.
Primarily office/remote knowledge-work role; prolonged periods of sitting and computer use. May require occasional off-hours support for critical incidents or maintenance windows. Travel requirements are minimal (0-10%) as applicable.
About ZoetisAt Zoetis, our purpose is to nurture the world and humankind by advancing care for animals. As a Fortune 500 company and the world leader in animal health, we discover, develop, manufacture and commercialize vaccines, medicines, diagnostics and other technologies for companion animals and livestock. We know our people drive our success. Our award-winning culture, built around our Core Beliefs, focuses on our colleagues' careers, connection…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).