Senior HPC DevOps Engineer
Job in
Rockville, Montgomery County, Maryland, 20849, USA
Listed on 2026-08-02
Listing for:
Peraton
Full Time
position Listed on 2026-08-02
Job specializations:
-
IT/Tech
SRE/Site Reliability, Unix/Linux, IT Infrastructure, Systems Engineer
Job Description & How to Apply Below
Required qualifications
- 12+ years of experience and a BS in computer science, IT, or related technical field, MS and 10 years of experience, or a Ph.D. with 8 years of experience. Four years of additional experience is required in lieu of a Bachelors’ degree for a total of 16 years of experience.
- 7+ years in Linux systems / SRE / Dev Ops, including production cluster operations in an HPC or large‑scale compute environment.
- 3+ years of experience building and operating Ansible automation at scale (roles/collections, idempotency, inventories, secrets).
- Strong Linux hardening & compliance fundamentals (SELinux/App Armor, SSH key automation, baseline config management).
- Demonstrated experience operating or automating clustered compute environments (HPC, large Linux farms, or similar).
- Hands‑on experience with container tooling in Linux environments, including image lifecycle/versioning.
- Familiarity with incident response and runbook‑driven operations; ability to automate common remediations.
- Strong Git workflow and documentation practices.
- Must hold at least one active/current technical certification from the following:
- Systems engineering (e.g., INCOSE)
- Information security (e.g., CISSP)
- Networking (e.g., CCNA)
- System Administration (e.g., RHCE, MCSE)
- Virtualization (e.g., VCP)
- IT systems management (e.g., ITIL)
- Project management (e.g., PMP, Agile)
- Active TS/SCI security clearance with a current polygraph is required.
- Bare‑metal provisioning experience (PXE/iPXE, Kickstart/Preseed, Foreman/MAAS) and hardware OOB management.
- CI/testing for automation and promotion pipelines for playbooks.
- Experience with tuned performance profiles, HPC performance troubleshooting, and GPU node health validation.
Peraton Labs is seeking a poly cleared Senior HPC Dev Ops Engineer to own the operations and automation lifecycle for an existing HPC/AI compute cluster (Linux). You will work closely with Peraton team members, as well as directly with our Maryland-based customer, in a fast‑paced environment at a customer site. In this role you will codify repeatable operations in Ansible and drive execution through an enterprise automation controller to enforce desired state, detect drift, accelerate node onboarding, and streamline incident response via runbook automation integrated with monitoring and ITSM.
This position requires full‑time on‑site work at a customer site near College Park, MD.
Key responsibilities- Automation ownership:
Own and manage automation workflows, including job templates, inventories, credentials, RBAC configurations, execution environments, and promotion across environments. - Desired‑state and drift detection:
Enforce desired state across cluster services via code‑driven configuration; implement drift detection and alert on deviations; reconcile runtime state vs configured state. - Compute node onboarding (Bare‑metal/VM):
Build and maintain an automated node bootstrap workflow that installs/configures the OS, applies security and performance baselines, enrolls nodes into the scheduler and shared storage ecosystem, validates hardware and service readiness (CPU, network, accelerator, storage mounts), and reports pass/fail results. - Patch & vulnerability response:
Implement rolling maintenance and patch automation to meet defined vulnerability response SLAs. Maintain version‑controlled container build definitions and integrate image scanning into the build/release lifecycle. - Logging & observability:
Ensure automation and operational workflows emit auditable logs to centralized analytics and integrate with metrics/alerting to enable reliable incident response, proactive detection, and safe auto‑remediation. - Incident/problem management:
Automate responses to common incidents (hung nodes, storage performance alarms, image vulnerabilities, hardware failures) leveraging out‑of‑band hardware management interfaces and standardized runbooks. - Docs‑as‑code:
Keep runbooks and operational documentation versioned alongside automation and publish operator guidance to the org’s documentation platform.
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×