×
Register Here to Apply for Jobs or Post Jobs. X

HPC Platform Product Lead

Job in Kalamazoo, Kalamazoo County, Michigan, 49006, USA
Listing for: Zoetis Spain SL
Full Time position
Listed on 2026-07-24
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, AI Business & Operations
Salary/Wage Range or Industry Benchmark: 130000 - 180000 USD Yearly USD 130000.00 180000.00 YEAR
Job Description & How to Apply Below
## HPC Platform Product Lead Apply locations:
Kalamazoo - Portage Roadtime type:
Full time posted on:
Posted Todayjob requisition :
JR
** POSITION SUMMARY
** Owns the end-to-end success of Zoetis’ VMRD High Performance Computing (HPC) environment as both a product and a platform: sets vision and roadmap, defines service offerings and governance, and provides technical leadership for architecture, SLURM scheduling, reliability, and sustainable growth. Partners with scientific/analytics users and business leaders to translate workload demand into a prioritized backlog and measurable value narrative. Ensures the environment delivers the right performance, cost, and risk posture through capacity planning, operational excellence, and data lifecycle strategy.##

POSITION RESPONSIBILITIES
** Product ownership & value management (25%)
*** Own product vision, roadmap, service catalog, and backlog with clear acceptance criteria.
* Translate utilization/throughput/outcomes into investment asks, funding recommendations, and executive-ready value stories (showback/chargeback as applicable).
** HPC architecture & technical standards (20%)
*** Define target-state architecture and standards across compute, storage, network, security, and tooling; drive multi-year evolution planning.
** SLURM scheduling leadership & L3 support (15%)
*** Own scheduling strategy (partitions/queues, QoS, fairshare, accounting, reservations); lead configuration, troubleshooting, and job efficiency analysis; provide technical backup as needed.
** Operations, reliability, and observability (20%)
*** Drive operational excellence via monitoring/alerting, dashboards, SLO/SLI practices, incident/problem follow-up, runbooks, and continuous improvement.
** Capacity planning, procurement, and lifecycle management (10%)
*** Plan and execute compute and storage growth; manage allocations; coordinate procurement and lifecycle refresh.
** Governance, access, and user enablement (10%)
*** Establish onboarding/entitlements/RBAC and auditability; improve user support processes (ticketing, escalations, knowledge base) and publish workload best practices (CPU/GPU optimization, execution standards, data movement).## EDUCATION AND EXPERIENCE
*
* Required:

*** Bachelor’s degree in computer science, Engineering, Information Systems, or equivalent practical experience.
* 7+ years of experience in infrastructure engineering/architecture, platform engineering, or adjacent technical product roles.
* 5+ years of experience in HPC environments (on-prem, cloud, and/or hybrid).
* Demonstrated experience partnering with scientists/engineers to translate workload needs into prioritized platform capabilities and measurable value stories.
** Preferred:
*** Master’s degree in a relevant field.
* Experience operating HPC in regulated or highly governed environments.
* Product Owner / Agile delivery experience (e.g., backlog management, acceptance criteria, road mapping).## TECHNICAL SKILLS REQUIREMENTS
* Strong SLURM administration skills: partitions/queues, QoS, fairshare, accounting, reservations; job monitoring and troubleshooting; job efficiency analysis; user guidance.
* HPC infrastructure knowledge:
Linux administration fundamentals, compute (CPU/GPU), high-speed networking, shared/parallel storage concepts, and capacity/performance planning.
* Automation/scripting (e.g., Bash/Python) and operational tooling; infrastructure-as-code/automation tooling where applicable.
* Observability: utilization telemetry, job analytics, dashboards/alerting for node health and storage utilization; actionable operational reporting.
* Data lifecycle and protection: retention policies, tiered storage, archive/backup strategy, restore testing, and operational runbooks.
* Product and delivery skills: requirements discovery, value storytelling, roadmap/backlog prioritization, stakeholder management, and communication of technical trade-offs and ROI.## PHYSICAL POSITION REQUIREMENTSPrimarily office/remote knowledge-work role; prolonged periods of sitting and computer use. May require occasional off-hours support for critical incidents or maintenance windows. Travel requirements are minimal (0–10%) as applicable.

Full time

Regular Colleague
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary