HPC Assistant Director
Listed on 2026-09-11
-
IT/Tech
Systems Engineer, IT Project Manager
Salary will be commensurate with the selected candidate's qualifications, experience, and education.
Job SummaryInformation and Technology Services (ITS) at the University of Michigan has an exciting opportunity for an HPC Assistant Director to lead the people, services, and systems that support large-scale research computing across the University, including regulated and researcher-owned environments. This is primarily a managerial and service-leadership role, accountable for team development, strategic and operational planning, and the availability, reliability, security, performance, capacity, and lifecycle of HPC services.
The Assistant Director will work in close partnership with the HPC Technical Lead, who serves as the principal technical advisor and leads HPC architecture and technical direction. The successful candidate does not need to be the team's foremost technical expert but must have sufficient understanding of Linux, HPC clusters, networking, and storage to evaluate recommendations, ask informed questions, weigh risks and tradeoffs, and make sound operational and organizational decisions.
This position also leads major incident response, establishes operational priorities and policies, defines requirements for significant procurements, and builds partnerships across ITS and the University. The successful candidate will bring a strong commitment to developing staff and enabling research, scholarship, and discovery.
The HPC Assistant Director will report to the Director of ITS-Advanced Research Computing.
Responsibilities*You will provide strategic and operational leadership for the team responsible for designing, building, and operating the University's high-performance computing (HPC) environments.
Key responsibilities include:
- Collaborate with the broader Advanced Research Computing (ARC) leadership team to develop and execute University-wide research computing strategies, service roadmaps, and investment priorities.
- Lead, manage, and develop a team of HPC and Linux professionals by setting clear expectations, providing coaching and feedback, supporting professional growth, and fostering an inclusive, collaborative, and accountable team culture.
- Plan and prioritize the team's work across infrastructure projects, system maintenance, security, technical debt, and operational support; allocate resources, manage dependencies and risks, and ensure reliable delivery aligned with ARC strategy and researcher needs.
- Develop and implement operational policies, standards, automation, metrics, and service-management practices that improve reliability, sustainability, user experience, and efficient use of University resources. Use meaningful service and capacity metrics to inform operational decisions and long-term investments.
- Own the operational health and lifecycle of ARC's HPC services, establishing expectations for availability, reliability, security, performance, capacity, supportability, observability, continuity, and disaster recovery.
- Lead the response to significant service disruptions, security events, and other production issues by coordinating technical teams, establishing priorities, assessing risk and service impact, communicating with stakeholders, and ensuring remediation and follow-through in accordance with the ITS major incident process.
- Engage researchers, other ITS units (e.g. data centers and networking), unit IT organizations, and the broader research IT community to understand evolving needs, translate requirements into technical and service capabilities, and establish clear expectations about feasible solutions.
- Develop sustainable financial and service-delivery models for HPC services, including cost projections, capacity planning, recharge rates, lifecycle…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).