×
Register Here to Apply for Jobs or Post Jobs. X

HPC Platform Product Lead

Job in Warren, Macomb County, Michigan, 48091, USA
Listing for: Zoetis Spain SL
Full Time position
Listed on 2026-09-11
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below

Position Summary

Owns the end‑to‑end success of Zoetis’ VMRD High Performance Computing (HPC) environment as both a product and a platform. Sets vision and roadmap, defines service offerings and governance, and provides technical leadership for architecture, SLURM scheduling, reliability, and sustainable growth. Partners with scientific/analytics users and business leaders to translate workload demand into a prioritized backlog and measurable value narrative. Ensures the environment delivers the right performance, cost, and risk posture through capacity planning, operational excellence, and data lifecycle strategy.

Responsibilities
  • Product ownership & value management (25%):
    Own product vision, roadmap, service catalog, and backlog with clear acceptance criteria. Translate utilization, throughput, and outcomes into investment asks, funding recommendations, and executive‑ready value stories (showback/chargeback as applicable).
  • HPC architecture & technical standards (20%):
    Define target‑state architecture and standards across compute, storage, network, security, and tooling; drive multi‑year evolution planning.
  • SLURM scheduling leadership & L3 support (15%):
    Own scheduling strategy (partitions/queues, QoS, fairshare, accounting, reservations); lead configuration, troubleshooting, and job efficiency analysis; provide technical backup as needed.
  • Operations, reliability, and observability (20%):
    Drive operational excellence via monitoring/alerting, dashboards, SLO/SLI practices, incident/problem follow‑up, runbooks, and continuous improvement.
  • Capacity planning, procurement, and lifecycle management (10%):
    Plan and execute compute and storage growth; manage allocations; coordinate procurement and lifecycle refresh.
  • Governance, access, and user enablement (10%):
    Establish onboarding, entitlements, RBAC, and auditability; improve user support processes (ticketing, escalations, knowledge base); publish workload best practices (CPU/GPU optimization, execution standards, data movement).
Education & Experience

Required:

Bachelor’s degree in computer science, engineering, information systems, or equivalent practical experience. 7+ years of experience in infrastructure engineering/architecture, platform engineering, or adjacent technical product roles. 5+ years of experience in HPC environments (on‑prem, cloud, and/or hybrid). Demonstrated experience partnering with scientists/engineers to translate workload needs into prioritized platform capabilities and measurable value stories.

Preferred:
Master’s degree in a relevant field; experience operating HPC in regulated or highly governed environments; product owner/Agile delivery experience (e.g., backlog management, acceptance criteria, road‑mapping).

Technical Skills Requirements
  • Strong SLURM administration skills: partitions, queues, QoS, fairshare, accounting, reservations; job monitoring and troubleshooting; job efficiency analysis; user guidance.
  • HPC infrastructure knowledge:
    Linux fundamentals, compute (CPU/GPU), high‑speed networking, shared/parallel storage concepts, capacity/performance planning.
  • Automation/scripting:
    Bash/Python; infrastructure‑as‑code/automation tooling where applicable.
  • Observability: utilization telemetry, job analytics, dashboards/alerting for node health and storage utilization; actionable operational reporting.
  • Data lifecycle and protection: retention policies, tiered storage, archive/backup strategy, restore testing, and operational runbooks.
  • Product and delivery skills: requirements discovery, value storytelling, roadmap/backlog prioritization, stakeholder management, and communication of technical trade‑offs and ROI.
Physical

Position Requirements

Primarily an office/remote knowledge‑work role; prolonged periods of sitting and…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary