DevOps Engineer
Lemont, DuPage County, Illinois, 60439, USA
Listed on 2026-07-31
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer
The Argonne Leadership Computing Facility’s (ALCF) mission is to accelerate major scientific discoveries and engineering breakthroughs for humanity by designing and providing world-leading computing facilities in partnership with the computational science community. We help researchers solve some of the world’s largest and most complex problems with our unique combination of supercomputing resources and computational science expertise.
The ALCF seeks a Dev Ops engineer to support infrastructure within ALCF and in support of the Department of Energy's American Science Cloud (AmSC) project. AmSC is a secure, federated, and science-optimized cloud environment that integrates the DOE’s world-leading computing and experimental facilities, data resources, and high-performance networks.
The AmSC platform enables DOE scientists to create, access, and integrate AI-ready datasets, run scalable AI model training and inference on leadership-class systems, perform distributed simulations, control instruments, and move data efficiently across sites.
As a Dev Ops Engineer, you will work within the Data Services and Workflows team at ALCF, with significant interaction with the Infrastructure Services group of AmSC to support all activities on our multi-cloud central hub infrastructure, for development, staging, pre-production, and production environments. Other L2 science service teams are deploying services on top of the infrastructure that the Infrastructure team manages – e.g. data catalogs and repositories, at-scale HPC compute services, user interfaces and APIs, and intelligent operations (AI/MLOps).
Your primary job responsibilities will be to support the science teams by building foundational infrastructure and developing CI/CD pipelines to deploy services on that infrastructure.
- The service stack is primarily Kubernetes-based. Perform cluster administration and application deployment assistance to users.
- Build and maintain pipelines for deploying cloud infrastructure and science services.
- Manage and use image registries such as Harbor.
- Writing and updating automation for resource provisioning and CI/CD pipelines – e.g. Terraform, Git Ops, Python.
- Implement security controls as defined by Cybersecurity team (Dev Sec Ops ).
- Configure basic instrumentation for infrastructure and core services, to feed into monitoring and alerting systems.
- Provide primary operational support and engineering for production applications.
- Define and implement KPIs, processes and drive continuous improvement.
- Diagnose platform operational problems quickly and effectively.
- Deploy, manage, and operate managed Kubernetes clusters (Amazon EKS, Azure AKS, Google GKE, or equivalent), including node group lifecycle management, cluster upgrades, cloud-native networking integrations (load balancer controllers, CNI plugins), and multi-environment promotion across dev, staging, and production.
- Coordinate with vendors to resolve hardware and software problems.
which applies to employees regularly scheduled for some onsite and some remote days, with employees typically working more than 60% of their time remotely. Non-exempt employees should not be permitted to split their time between on-site and remote work on a given workday unless they have advance supervisor approval.
Position Requirements- PT2:
Bachelor’s Degree in computer science or closely related field and a minimum of 2+ years of experience as a Dev Ops engineer and/or Cloud Engineer. - US citizenship:
To perform the essential functions of this position successful applicants must provide proof of U.S. citizenship, which is required to comply with federal regulations and contract - Previous experience as a team lead – able to perform task management for Dev Ops or Cloud Engineering teams.
- Excellent interpersonal/communication skills, and the ability to work as part of a team.
- Working knowledge of cloud application architecture patterns and a thorough grasp of common products and managed services for at least one Cloud Service Provider (e.g. AWS).
- Working knowledge of Kubernetes cluster administration and concepts (CR/CRDs) and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).