Computing; HPC) Engineer
Listed on 2026-09-12
-
IT/Tech
Systems Engineer, Unix/Linux, Cloud Computing: Infrastructure & Operations
The Aerospace Corporation is the trusted partner to the nation’s space programs, solving the hardest problems and providing unmatched technical expertise. As the operator of a federally funded research and development center (FFRDC), we are broadly engaged across all aspects of space— delivering innovative solutions that span satellite, launch, ground, and cyber systems for defense, civil and commercial customers. When you join our team, you’ll be part of a special collection of problem solvers, thought leaders, and innovators.
Join us and take your place in space. The Aerospace Corporation is seeking a talented High-Performance Computing (HPC) Engineer (Site Reliability Engineer Staff III/IV) to join our Computational Services team. In this role, you will develop, implement, and optimize HPC clusters that support both on-premises and cloud environments. You will work alongside rocket scientists and engineers, tackling complex space enterprise challenges while having direct impact on critical national security missions.
We value a collaborative, proactive mindset and a shared commitment to engineering excellence. The selected candidate will be required to work full-time, on-site at our facility in El Segundo, CA or Chantilly, VA.
Collaborate with scientists and engineers on diverse projects supporting mission‑critical technical analysis for national space assets Lead cross‑functional teams and mentor junior engineers. Design and implement HPC solutions that optimize resource utilization across diverse workloads in both classified and unclassified settings. Manage on premise 10,000-core classified cluster and a 5,000-core unclassified cluster to ensure peak performance. Deliver high‑quality HPC infrastructure design, and system configuration.
Develop and deploy automation solutions using tools such as Clush. Manage infrastructure using Infrastructure‑as‑Code and Git Ops practices Implement, support, and optimize GPU computing. Monitor, analyze, and tune HPC system performance, utilization, and resource allocation to maintain operational efficiency. Develop cost‑efficient HPC service offerings that align with mission and business objectives. Harden Linux systems to meet stringent security requirements.
Minimum Requirements for the Site Reliability Engineer Staff III:
Bachelor’s degree in Computer Science, Engineering, or equivalent experience. Minimum of 7 years’ experience in Linux system administration within an enterprise HPC environment. Experience supporting technical software (compilers, mod&sim tools, languages, COTS, GOTs) including the development of environment modules. In‑depth knowledge of Linux, networking, and HPC systems. Experience with Infrastructure‑as‑Code and Git Ops Proven experience in managing the Slurm scheduler and setting up HPC systems for both interactive and batch workloads.
Experience provisioning and supporting AI & NVIDIA GPU technologies(e.g. CUDA) Proficiency in scripting and competence with automation tools such as Clush. Experience hardening Linux systems to meet security requirements Experience with hardware and infrastructure automation in environments using server vendors such as HPE or Cisco. Strong communication skills, with an ability to work both independently and as part of a geographically distributed team.
CompTIA Security+ CE certification or equivalent that meets DoD 8570.01‑m requirements for IAT Level II personnel Ability to obtain and maintain a TS/SCI clearance (U.S. citizenship required). Demonstrated ability to lead cross‑functional teams and mentor junior engineers.
In addition to the above, the minimum requirements for the Site Reliability Engineer Staff IV include: 9+ years of…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).