AI Research Computing Infrastructure Engineer
Listed on 2026-07-11
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations
AI Research Computing Infrastructure Engineer
Job Employee Type: exempt full-time Division:
Enterprise Information Technology Facility:
Frederick:
Ft Detrick
Location:
PO Box B, Frederick, MD 21702 USA
The Frederick National Laboratory is operated by Leidos Biomedical Research, Inc. The lab addresses some of the most urgent and intractable problems in the biomedical sciences in cancer and AIDS, drug development and first-in-human clinical trials, applications of nanotechnology in medicine, and rapid response to emerging threats of infectious diseases.
Accountability, Compassion, Collaboration, Dedication, Integrity and Versatility; it’s the FNL way.
PROGRAM DESCRIPTIONThe mission of Enterprise Information Technology (EIT) is to develop an enterprise-level, consolidated information technology infrastructure that provides exceptional IT capabilities to the Frederick National Labs for Cancer Research (NCI-Frederick/FNLCR) in support of basic, translational, and clinical cancer and AIDS research. The IT Operations Group (ITOG) is a part of Enterprise Information Technology (EIT) within Leidos Biomedical Research, Inc. ITOG is responsible for computational servers, storage servers, virtual machine infrastructure, and the FNLCR network.
ITOG focuses on implementing enterprise IT best practices in the areas of computational services, storage, backup, and archiving; batch and application support; server consolidation and virtualization; network infrastructure; unification of voice, teleconferencing, and video communication technologies; and improved infrastructure for collocation of dedicated servers.
The Research Computing Infrastructure Engineer will design, build, and operate next-generation high-performance computing (HPC) environments that support container-based workflows and GPU-accelerated research computing. The position will play a key role in evaluating, implementing, and maintaining scalable and secure computing architectures for advanced data analysis, AI/ML model training, and simulation workloads. The engineer will collaborate closely with researchers, IT professionals, and external partners to translate scientific requirements into reliable, high-performance computing solutions.
- Design and implement next-generation high-performance computing (HPC) environments that leverage container-driven workflows for GPU-accelerated research.
- Build and maintain container orchestration systems for batch and distributed workloads.
- Integrate containerized job workflows with existing HPC schedulers and storage systems.
- Develop and maintain job templates for batch GPU training and multi-node distributed computing.
- Automate deployment, configuration, and scaling through infrastructure-as-code and CI/CD practices.
- Monitor, benchmark, and optimize system performance, reliability, and resource utilization.
- Collaborate with researchers to containerize and optimize legacy workflows for scalable execution.
- Lead evaluation of emerging tools (e.g., Prefect, Ray, Airflow, Dagster) for workflow orchestration and distributed computing.
- Contribute to the development of tools and bridges between orchestration frameworks and traditional HPC environments.
To be considered for this position, you must minimally meet the knowledge, skills, and abilities listed below:
- Possession of Bachelor’s degree from an accredited college/university according to the Council for Higher Education Accreditation (CHEA) or four (4) years relevant experience in lieu of degree. Foreign degrees must be evaluated for U.S. equivalency.
- In addition to the education requirement, a minimum of eight (8) years of related experience.
- Strong Linux systems engineering and administration experience.
- Hands-on experience with container orchestration tools such as Kubernetes, Nomad, Run:
AI, etc. - Hands-on experience with scripting/programming skills (Python, Bash, or Go) for automation, monitoring, and job orchestration.
- Experience with infrastructure-as-code / automation tooling (Terraform, Ansible, Packer, or equivalent).
- Familiarity with system performance analysis, monitoring, and tuning.
- Comfortable with small-team…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).