Sr. HPC System Administrator
Listed on 2026-09-12
-
IT/Tech
Systems Administrator, IT Support, Systems Engineer, Network Administrator
Department Provost Research Computing Center About the Department
The University of Chicago Research Computing Center (RCC), a unit in the Office of Research, provides high-end research computing resources to researchers at the University of Chicago. It is dedicated to enabling research by providing access to centrally managed High-Performance Computing (HPC), storage, and visualization resources. These resources include hardware, software, high-level scientific and technical user support, and the education and training required to help researchers make full use of modern HPC technology and local and national supercomputing resources.
The Office of Research oversees the conduct of sponsored research, research program development, and contract management functions.
The job uses specialized knowledge and breadth of expertise to design automated, scalable, and rapidly deployable solutions to infrastructure development and server configuration. Leads installation, configuration, and maintenance of operating systems. Uses best practices and systems knowledge to monitor and alert systems, utility software, and firewalls. Guides maintenance for production servers as well as Windows and Linux servers. The University of Chicago is seeking a highly qualified Senior HPC System Administrator to join the system and operation team that builds and manages RCC HPC systems and facility operations.
The individual in this position will be involved in the procurement and management of HPC hardware and software. This is a hybrid position requiring 3 days onsite.
The University of Chicago uses AI-assisted tools to streamline and augment some recruitment processes; however, AI is not used to make hiring decisions.
Responsibilities- Installing, configuring, and maintaining large computer clusters/servers and software.
- Day-to-day operations of the systems including systems administration, monitoring and storage performance up to and including network components.
- Management of the system’s network switch, parallel file system and HPC software stack and tools.
- Configuration of the scheduling and queuing system.
- Diagnosing and resolving system operational problems quickly and effectively.
- Coordinating with vendors to resolve hardware and software problems.
- Assist users with access and other help desk ticket requests or issues.
- Use scripting/programming skills to enable system-level automation, problem detection, security maintenance and patch management.
- Building and deploying open-source software and software from vendors/partners.
- Providing reliable and efficient backups/restores for all managed systems.
- Documenting system administration procedures for routine and complex tasks.
- Maintaining and monitoring the security of the HPC systems and servers.
- Plans and installs necessary patches and upgrades for servers and their associated storage, network, communications, and peripheral sub-systems.
- Installs and maintains an appropriate level of intrusion detection, monitoring, and auditing software as required.
- Tracks compliance and maintains documentation for hardware, software, and service inventories for management reports.
- Performs other related work as needed.
Education:
Minimum requirements include a college or university degree in related field.
Work Experience:
Minimum requirements include knowledge and skills developed through 5-7 years of work experience in a related job discipline.
Certifications:
---
Education:
Master's degree in Computer Science or closely related field.
Experience:
Full time Linux system administration experience in a large distributed computing environment. Previous experience in providing support for Linux HPC cluster used for scientific research.…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).