×
Register Here to Apply for Jobs or Post Jobs. X

HPC Platform Product Lead

Job in Lansing, Ingham County, Michigan, 48900, USA
Listing for: Zoetis Spain SL
Full Time position
Listed on 2026-07-24
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, AI Business & Operations, SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below

Position Summary

Owns the end‑to‑end success of Zoetis’ VMRD High Performance Computing (HPC) environment as both a product and a platform. Sets vision and roadmap, defines service offerings and governance, and provides technical leadership for architecture, SLURM scheduling, reliability, and sustainable growth. Partners with scientific/analytics users and business leaders to translate workload demand into a prioritized backlog and measurable value narrative. Ensures the environment delivers the right performance, cost, and risk posture through capacity planning, operational excellence, and data lifecycle strategy.

Responsibilities
  • Product ownership & value management (25%):
    Own product vision, roadmap, service catalog, and backlog with clear acceptance criteria. Translate utilization, throughput, and outcomes into investment asks, funding recommendations, and executive‑ready value stories (showback/chargeback as applicable).
  • HPC architecture & technical standards (20%):
    Define target‑state architecture and standards across compute, storage, network, security, and tooling; drive multi‑year evolution planning.
  • SLURM scheduling leadership & L3 support (15%):
    Own scheduling strategy (partitions/queues, QoS, fairshare, accounting, reservations); lead configuration, troubleshooting, and job efficiency analysis; provide technical backup as needed.
  • Operations, reliability, and observability (20%):
    Drive operational excellence via monitoring/alerting, dashboards, SLO/SLI practices, incident/problem follow‑up, runbooks, and continuous improvement.
  • Capacity planning, procurement, and lifecycle management (10%):
    Plan and execute compute and storage growth; manage allocations; coordinate procurement and lifecycle refresh.
  • Governance, access, and user enablement (10%):
    Establish onboarding, entitlements, RBAC, and auditability; improve user support processes (ticketing, escalations, knowledge base); publish workload best practices (CPU/GPU optimization, execution standards, data movement).
Education & Experience

Required:

Bachelor’s degree in computer science, engineering, information systems, or equivalent practical experience. 7+ years of experience in infrastructure engineering/architecture, platform engineering, or adjacent technical product roles. 5+ years of experience in HPC environments (on‑prem, cloud, and/or hybrid). Demonstrated experience partnering with scientists/engineers to translate workload needs into prioritized platform capabilities and measurable value stories.

Preferred:
Master’s degree in a relevant field; experience operating HPC in regulated or highly governed environments; product owner/Agile delivery experience (e.g., backlog management, acceptance criteria, road‑mapping).

Technical Skills Requirements
  • Strong SLURM administration skills: partitions, queues, QoS, fairshare, accounting, reservations; job monitoring and troubleshooting; job efficiency analysis; user guidance.
  • HPC infrastructure knowledge:
    Linux fundamentals, compute (CPU/GPU), high‑speed networking, shared/parallel storage concepts, capacity/performance planning.
  • Automation/scripting:
    Bash/Python; infrastructure‑as‑code/automation tooling where applicable.
  • Observability: utilization telemetry, job analytics, dashboards/alerting for node health and storage utilization; actionable operational reporting.
  • Data lifecycle and protection: retention policies, tiered storage, archive/backup strategy, restore testing, and operational runbooks.
  • Product and delivery skills: requirements discovery, value storytelling, roadmap/backlog prioritization, stakeholder management, and communication of technical trade‑offs and ROI.
Physical

Position Requirements

Primarily an office/remote knowledge‑work role; prolonged periods of sitting and computer use. May require occasional off‑hours support for critical incidents or maintenance windows. Travel requirements are minimal (0–10%) as applicable.

Legal & EEO Statement

Zoetis is committed to equal opportunity in the terms and conditions of employment for all employees and job applicants without regard to race, color, religion, sex, sexual orientation, age, gender identity or expression, national origin, disability, veteran status, or any other protected classification. Disabled individuals are given an equal opportunity to use our online application system. We offer reasonable accommodations as an alternative if requested by an individual with a disability.

All applicants must possess or obtain authorization to work in the U.S. for Zoetis. Zoetis retains sole and exclusive discretion to pursue sponsorship for the acquisition or maintenance of non‑immigrant status and employment eligibility, considering factors such as availability of qualified U.S. workers. Individuals requiring sponsorship must disclose this fact.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary