HPC DevOps Engineer
Listed on 2026-07-27
-
IT/Tech
Systems Engineer, IT Support, Cloud Computing: Infrastructure & Operations
One of the best college towns and best places to live in the United States, Chapel Hill has diverse social, cultural, recreation and professional opportunities that span the campus and community.
University employees can choose from a wide range of professional training opportunities for career growth, skill development and lifelong learning and enjoy exclusive perks for numerous retail, restaurant and performing arts discounts, savings on local child care centers and special rates on select campus events. UNC
-Chapel Hill offers full-time employees a comprehensive benefits package , paid leave, and a variety of health, life and retirement plans and additional programs that support a healthy work/life balance.
ITS Research Computing (RC) aims to provide a world-class computing infrastructure as well as other technology tools and capabilities to support the research needs of University faculty and staff. Its goal is to provide a state-of-the-art environment to support the highest level of multidisciplinary research and help UNC
-Chapel Hill become the premier research university in the United States.
This position may be eligible for a hybrid work arrangement that may include a partially remote work location, consistent with System Office policy. UNC Chapel Hill employees are generally required to reside within a reasonable commuting distance of their assigned duty station.
The HPC Dev Ops Engineer will have a broad role within Research Computing at UNC Chapel Hill, designing, deploying, operating and maintaining services and solutions in support of leading
-edge academic research, primarily related to High Performance Computing ( HPC ) as well as High Throughput Computing ( HTC ).
Responsibilities of the HPC Dev Ops Engineer include leading and contributing to a variety of areas within Research Computing related to the design, build, and operation of HTC / HPC services along with ancillary systems including high-speed storage and networking, containers, databases, monitoring, orchestration services, etc.
The ideal candidate for this position should possess a comprehensive understanding of Linux systems administration in HPC environments, broad technical capabilities, and enjoy applying these skills in academic research to further the mission of the University.
The successful candidate should possess excellent leadership and communication skills to effectively lead complex projects. They should have the ability to engage directly with faculty and researchers, as well as leverage research computing communities of practice beyond Carolina.
ITS Research Computing aims to provide world-class computing and data infrastructure as well as other services and capabilities in support of research needs for faculty, staff, students, and collaborators. Our goal is to provide state
-of-the-art environments and services supporting the highest level of multidisciplinary research.
Master’s and 1-2 years’ experience; or Bachelors and 2-4 years’ experience; or will accept a combination of related education and experience in substitution.
Required Qualifications , Competencies, and Experience- Proficiency with Linux operating systems, including installation, configuration, troubleshooting, and lifecycle management in an enterprise or research-computing context. Experience with automation and configuration-management tools such as Cobbler, SALT , or comparable platforms used to deploy and manage systems iliarity with enterprise storage platforms, including routine configuration tasks, capacity monitoring, health assessment, and basic hardware maintenance such as drive replacement. Ability to interpret and act on monitoring, logging, and alerting data to maintain operational continuity across diverse service lines.
Understanding of RHEL repo management and practices for maintaining consistent, secure OS-layer operations. - Ability to lift and maneuver equipment up to 50 pounds with or without reasonable accommodations, work in hot/cold aisle conditions, and perform tasks in confined rack environments. Capacity to stand, bend, and work in data-center spaces for extended periods while performing installation, cabling, and hardware maintenance tasks.
- Strong diagnostic and problem-solving abilities, including the capacity to assess hardware and system issues under time constraints. Demonstrated adherence to change-management and operational best practices in complex technical environments. Ability to coordinate effectively with vendors, service providers, and internal stakeholders to support hardware and storage operations.
- Clear written and verbal communication skills, including the ability to convey technical information to both technical and non-technical audiences. Ability to work collaboratively within a team where some responsibilities are shared and others are independently owned.
- Understanding of operational practices required to support…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).