Computing; HPC) Systems Engineer
Listed on 2026-09-12
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations
At Jefferson Lab,you'llchampioncutting-edge science and operational excellence while shaping the future of discovery. Join us and make your mark - where excellence meets purpose, andgreat minds truly
matter.
The good-faith pay range for this role is $91,800 - $145,050 per year. Actual compensation may vary and may be above the posted range based on factors such as a candidate's skills, experience, education, certifications, and work location.
What your job will be like:We are seeking a High-Performance Computing (HPC) Systems Engineer to architect, deploy, and maintain the large-scale physical hardware, distributed file systems, and low-latency networking infrastructure powering our scientific computing ecosystem. This role focuses on the bare-metal and system-level foundations of a petabyte-scale environment, ensuring high availability, peak storage performance, and reliable data movement for experimental nuclear physics workloads. The ideal candidate will blend deep Linux systems administration expertise with modern infrastructure-as-code automation to support state-of-the-art research computing clusters.
Inthis job you will:
- Collaboratively design and implement software infrastructure supporting of High Performance and High Throughput computing using best-in-class containerization and orchestration tools.
- Deploy, maintain, and operate physical infrastructure for scientific computing services.
- Create highly available services through consideration of the entire stack from hardware and networking though the user application.
- Engage with users to meet service requirements for performance and availability.
- Develop software to address gaps in existing tools.
- Audit and recommend architectural changes to improve performance, security, and availability.
- Contribute to upstream development efforts of community tools.
- Work with students and interns on research projects.
- Required:
3 or more years Experience performing enterprise linux System Administration tasks including installation, configuration, and support of COTS/GOTS/FOSS software, file, network, and large-scale storage systems. - Required:
Experience operating in a production environment with high availability requirements. - Required:
Experience with automating test procedures - Preferred:
Experience with NP/HEP HPC infrastructure like slurm, Rucio, Globus, XrootD
- Required:
Bachelor's Degree in Computer Science, Computer Engineering, or related degree with significant Computer Science coursework - Preferred:
Master's Degree Computer Science, Data Science, or related discipline
Education above the minimum may be substituted for experience. Relevant experience may not be substituted for education.
Knowledge, Skills, and Abilities- Ability to architect, provision, tune, and maintain petabyte-scale, high-performance parallel and distributed storage systems, including Lustre, Ceph, CephFS, and NFS.
- Expert Kubernetes, Docker/Podman, and containerization administration.
- Expert Knowledge of Linux system administration including installation and configuration management (ie. Ansible/Puppet/Foreman)
- Ability to develop, debug, and test applications based upon design and performance requirements.
- Ability to communicate clearly in writing (e.g., email, presentations, drawings) and explain work to their supervisor and others in the group.
- Ability to work effectively with peers and participate in participate in troubleshooting, including off-hours during outages and as part of an occasional on-call rotation.
Join a community with a common purpose of solving the most challenging scientific and engineering problems of our time. The Jefferson Lab campusis located in southeastern
Virginia amidst a vibrant and growing technology community.
A career at Jefferson Lab is more than a job. You will be part of "big science" and work alongside top scientists and engineers from around the world unlocking the secrets of our visible universe. Managed by SURATech, LLC, Thomas Jefferson National Accelerator Facility is entering an exciting period of mission growth and is seeking new team members ready to apply their skills and passion to have an impact.
You could call it work, or you could call it a mission. We call it a challenge. We do things that will change the world.
- * Medical, Dental, and Vision Care Plans
* Flexible Spending Accounts - * Paid Time-off and Leave Programs (Paid Parental, vacation, holidays, and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).