Senior HPC Systems Administrator
Berkeley, Alameda County, California, 94709, USA
Listed on 2026-09-25
-
IT/Tech
Systems Administrator, Unix/Linux, IT Support, Systems Engineer
Red Line Performance Solutions (Red Line) has been in the HPC solutions engineering services business for 26 years and is consistently determined to keep the "bar of excellence" quite high for new hires. This enables Red Line to accomplish what other firms cannot and promotes a high level of staff retention. We offer services ranging from full life cycle HPC systems engineering to remote managed services to HPC program analysis.
Red Line is looking for a Senior High Performance Computing (HPC) Administrator to join our team. This position will work on various on-premise installations in support of new flagship supercomputer and managed infrastructure located at Lawrence Berkeley National Laboratory. The ideal candidate will be an experienced individual with a strong security, Linux, HPC, configuration management, systems automation and networking background. This position supports requires strong organization, management and customer facing skills.
The customer supports a high-visibility, next-generation national computing initiative supporting cutting-edge scientific research and discovery. As a Sr. HPC System Administrator, you will directly contribute to the reliability, performance, and availability of the advanced computing environment that enables large-scale scientific research and breakthroughs with far-reaching impact.
US citizenship and the ability to obtain a Public Trust clearance is a requirement to apply. This is a remote position but will work Pacific Time core hours. This full-time position offers a full benefits package including paid time off, 401k match, and health care benefits.
- paid time off
- 401k match
- health care benefits
Lead a team to administer resources from the operating system and above within on-premise HPC environments. Efforts include, but are not limited to:
- Integration and configuration of compute resources and service nodes
- All software installations on compute resources, services nodes, and parallel file systems
- Develop/implement system and performance monitoring and benchmarking
- Maintain system documentation
- Evaluate performance impacts of planned operating system changes
- Lead resource optimization and job scheduling software and policies
- Provide technical support to researchers using HPC resources, troubleshoot problems and develop appropriate computational strategies
- Provide emergency support on a 24x7 basis.
- Provide technical leadership and direction for other team members. Maintain team focus on production uptime and model performance.
- Review and present all change management requests to customer management.
- Work both independently and as part of the team; able to concurrently work on several projects.
- Effectively communicate with people of diverse backgrounds and computer knowledge.
- Manage individual and team task pipelines, ensuring all deliverables are met and tracking systems are consistently updated.
- Solicit and analyze customer feedback, ensuring critical issues are effectively communicated to the team and captured as trackable deliverables.
- Manage individual and team task pipelines, ensuring all deliverables are met and tracking systems are consistently updated.
- Solicit and analyze customer feedback, ensuring critical issues are effectively communicated to the team and captured as trackable deliverables.
- Travel to the client site in California approximately 1 week per quarter with some additional travel upon installation of the system.
- Minimum of 10 years Red Hat, Rocky, and/or CentOS Linux system administrator experience.
- Technical leadership experience in a large, production environment – leadership for the technical solution and the system administration team.
- Demonstrated ability to configure, deploy and manage major system areas such as batch system, network, data storage, backup system, database system, or distributed computing
- Experience with configuration management tools (e.g., Ansible)
- Ability to work both independently and as part of the team; flexibility in dealing with assignments and in working on several projects simultaneously
- Ability to effectively communicate with people of diverse backgrounds and computer knowledge.
- HPC system administration experience is highly preferred.
- Experience with batch systems (e.g., SLURM)
- Experience managing parallel and cluster file systems (e.g., Lustre)
- Network management experience, including in an HPC context (e.g., Infini Band)
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).