HPC Infrastructure & Cluster Engineer
Listed on 2026-09-09
-
IT/Tech
Systems Engineer, IT Infrastructure, Systems Administrator, Cloud Computing: Infrastructure & Operations
** ACTIVE TS/SCI SECURITY CLEARANCE REQUIRED**
We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environment. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.
Key Responsibilities:
- Cluster Administration:
Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades - Resource and Job Management:
Configure, maintain, and optimize workload management and orchestration platforms, utilizing Run:
AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster - Infrastructure Optimization:
Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads - Storage and Network Management:
Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an Infini Band GPU-to-GPU network infrastructure to minimize latency for distributed operations - Environment Configuration:
Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat Open Shift, required for seamless customer model deployment - Security and Compliance:
Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations
Basic Qualifications
- 5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments
- Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with Infini Band)
- Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g. Run:AI, SLURM)
- Hands‑on experience with enterprise container orchestration platforms, specifically Open Shift or Kubernetes
- Experience writing automation and configuration scripts (e.g. Bash, Python) to streamline cluster maintenance
- Proven ability to diagnose and resolve complex hardware, network, and OS-level issues
Preferred Qualifications
- Familiarity with parallel file systems and high-throughput storage architecture
- Prior experience engineering or managing high-speed GPU-to-GPU communication topologies
Additional Information
- All your information will be kept confidential according to EEO guidelines.
- Compensation is unique to each candidate and relative to the skills and experience they bring to the position. The salary range for this position is typically $170-$180k. This does not guarantee a specific salary as compensation is based upon multiple factors such as education, experience, certifications, and other requirements, and may fall outside of the above-stated range.
- Highlights of our benefits include Health/Dental/Vision, 401(k) match, Accrued PTO, STD/LTD/Life Insurance, Referral Bonuses, professional development reimbursement, and more!
D2 Technical Services is committed to a merit-based recruitment process and encourages applications from all qualified individuals. As a Veteran‑Owned Small Business, we particularly welcome applications from veterans who have the requisite skills and experience. Job applicants that are interested in one of our openings and may require a reasonable accommodation to participate in the job application or interview process, should contact us to request an accommodation.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).