Research Computing Infrastructure Engineer
Listed on 2026-07-22
-
IT/Tech
Systems Engineer, IT Infrastructure, Cloud Computing: Infrastructure & Operations
Research Computing Infrastructure Engineer
Job
Employee Type: exempt full-time
Division:
Enterprise Information Technology
Facility:
Frederick:
Ft Detrick
Location:
PO Box B, Frederick, MD 21702 USA
The Frederick National Laboratory is operated by Leidos Biomedical Research, Inc. The lab addresses some of the most urgent and intractable problems in the biomedical sciences in cancer and AIDS, drug development and first-in-human clinical trials, applications of nanotechnology in medicine, and rapid response to emerging threats of infectious diseases.
Accountability, Compassion, Collaboration, Dedication, Integrity and Versatility; it's the FNL way. PROGRAM DESCRIPTIONThe mission of Enterprise Information Technology (EIT) is to develop an enterprise-level, consolidated information technology infrastructure that provides exceptional IT capabilities to the Frederick National Labs for Cancer Research (NCI-Frederick/FNLCR) in support of basic, translational, and clinical cancer and AIDS research. The IT Operations Group (ITOG) is a part of Enterprise Information Technology (EIT) within Leidos Biomedical Research, Inc. ITOG is responsible for computational servers, storage servers, virtual machine infrastructure, and the FNLCR network.
ITOG focuses on implementing enterprise IT best practices in the areas of computational services, storage, backup, and archiving; batch and application support; server consolidation and virtualization; network infrastructure; unification of voice, teleconferencing, and video communication technologies; and improved infrastructure for collocation of dedicated servers.
The Research Computing Infrastructure Engineer provides technical leadership for the design, standards, and implementation of the SOM virtualization and container platform and its supporting infrastructure. This role evaluates and implements high-impact solutions involving current platform technologies, weighing long-term service delivery, cost, and operational practicality, and serves as a technical mentor across the team. The position supports research computing programs and operations within a federal compliance boundary.
- Design, deploy, and operate the SOM virtualization and container platform as secure, compliant infrastructure within a federal boundary.
- Operate the virtualization substrate: cluster lifecycle, node management, upgrades, and the underlying storage and networking.
- Manage downstream Kubernetes cluster provisioning, RBAC, and multi-tenant access for research groups.
- Own storage integration across the platform (software-defined storage plus POSIX and object storage such as VAST) for VM and container workloads.
- Build and maintain the security and compliance posture of the platform: SSP development, ATO support, continuous monitoring, and remediation under NIST 800-53/FISMA.
- Automate provisioning, configuration, and scaling through infrastructure-as-code and CI/CD practices (Ansible, Terraform, Packer, Git Hub Actions, or equivalent).
- Partner with embedded bioinformatics staff who own user-facing VM provisioning, container workflows, and job templates, providing the platform and guardrails they build on.
- Possession of Bachelor’s degree from an accredited college/university according to the Council for Higher Education Accreditation (CHEA) or four (4) years relevant experience in lieu of degree. Foreign degrees must be evaluated for U.S. equivalency.
- In addition to the education requirement, a minimum of eight (8) years of related experience. with strong Linux systems engineering and administration.
- Hands-on experience operating a virtualization platform in production (KVM/libvirt, VMware/vSphere, Open Stack, Harvester, or equivalent), including host lifecycle, live migration, and storage backends.
- Hands-on scripting/programming proficiency (Python, Bash, or Go) for automation and operations.
- Experience with infrastructure-as-code and automation tooling (Ansible, Terraform, Packer, or equivalent).
- Hands-on experience standing up and maintaining Kubernetes…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).