More jobs:
HPC Platform Product Lead
Job in
Kalamazoo, Kalamazoo County, Michigan, 49006, USA
Listed on 2026-07-24
Listing for:
Zoetis Spain SL
Full Time
position Listed on 2026-07-24
Job specializations:
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AI Business & Operations
Job Description & How to Apply Below
Kalamazoo - Portage Roadtime type:
Full time posted on:
Posted Todayjob requisition :
JR
** POSITION SUMMARY
** Owns the end-to-end success of Zoetis’ VMRD High Performance Computing (HPC) environment as both a product and a platform: sets vision and roadmap, defines service offerings and governance, and provides technical leadership for architecture, SLURM scheduling, reliability, and sustainable growth. Partners with scientific/analytics users and business leaders to translate workload demand into a prioritized backlog and measurable value narrative. Ensures the environment delivers the right performance, cost, and risk posture through capacity planning, operational excellence, and data lifecycle strategy.##
POSITION RESPONSIBILITIES
** Product ownership & value management (25%)
*** Own product vision, roadmap, service catalog, and backlog with clear acceptance criteria.
* Translate utilization/throughput/outcomes into investment asks, funding recommendations, and executive-ready value stories (showback/chargeback as applicable).
** HPC architecture & technical standards (20%)
*** Define target-state architecture and standards across compute, storage, network, security, and tooling; drive multi-year evolution planning.
** SLURM scheduling leadership & L3 support (15%)
*** Own scheduling strategy (partitions/queues, QoS, fairshare, accounting, reservations); lead configuration, troubleshooting, and job efficiency analysis; provide technical backup as needed.
** Operations, reliability, and observability (20%)
*** Drive operational excellence via monitoring/alerting, dashboards, SLO/SLI practices, incident/problem follow-up, runbooks, and continuous improvement.
** Capacity planning, procurement, and lifecycle management (10%)
*** Plan and execute compute and storage growth; manage allocations; coordinate procurement and lifecycle refresh.
** Governance, access, and user enablement (10%)
*** Establish onboarding/entitlements/RBAC and auditability; improve user support processes (ticketing, escalations, knowledge base) and publish workload best practices (CPU/GPU optimization, execution standards, data movement).## EDUCATION AND EXPERIENCE
*
* Required:
*** Bachelor’s degree in computer science, Engineering, Information Systems, or equivalent practical experience.
* 7+ years of experience in infrastructure engineering/architecture, platform engineering, or adjacent technical product roles.
* 5+ years of experience in HPC environments (on-prem, cloud, and/or hybrid).
* Demonstrated experience partnering with scientists/engineers to translate workload needs into prioritized platform capabilities and measurable value stories.
** Preferred:
*** Master’s degree in a relevant field.
* Experience operating HPC in regulated or highly governed environments.
* Product Owner / Agile delivery experience (e.g., backlog management, acceptance criteria, road mapping).## TECHNICAL SKILLS REQUIREMENTS
* Strong SLURM administration skills: partitions/queues, QoS, fairshare, accounting, reservations; job monitoring and troubleshooting; job efficiency analysis; user guidance.
* HPC infrastructure knowledge:
Linux administration fundamentals, compute (CPU/GPU), high-speed networking, shared/parallel storage concepts, and capacity/performance planning.
* Automation/scripting (e.g., Bash/Python) and operational tooling; infrastructure-as-code/automation tooling where applicable.
* Observability: utilization telemetry, job analytics, dashboards/alerting for node health and storage utilization; actionable operational reporting.
* Data lifecycle and protection: retention policies, tiered storage, archive/backup strategy, restore testing, and operational runbooks.
* Product and delivery skills: requirements discovery, value storytelling, roadmap/backlog prioritization, stakeholder management, and communication of technical trade-offs and ROI.## PHYSICAL POSITION REQUIREMENTSPrimarily office/remote knowledge-work role; prolonged periods of sitting and computer use. May require occasional off-hours support for critical incidents or maintenance windows. Travel requirements are minimal (0–10%) as applicable.
Full time
Regular Colleague
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×