×
Register Here to Apply for Jobs or Post Jobs. X

HPC Infrastructure & Cluster Engineer

Job in Springfield, Fairfax County, Virginia, 22161, USA
Listing for: teKnoluxion Consulting LLC
Full Time position
Listed on 2026-09-05
Job specializations:
  • IT/Tech
    Systems Engineer, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 148000 - 179000 USD Yearly USD 148000.00 179000.00 YEAR
Job Description & How to Apply Below

Overview HPC Infrastructure & Cluster Engineer
Springfield, VA

Active TS (SCI eligibility) clearance and eligibility to obtain a CI poly

At Bcore, our strength comes from how we deliver impact to the mission. Whether it’s architecting critical IT solutions, producing actionable intelligence, or developing cutting edge technology, we succeed because of the expertise, collaboration, and agility of our teams. Our Mission Services division combines enterprise IT, cloud solutions, Dev Sec Ops , systems engineering, software development, and operational support. Bcore accelerates decisive advantage for warfighters and intelligence professionals by fusing human insight, rapid-fire engineering, precision-measured outcomes, and relentless grit into mission-ready solutions.

Do you want to join a team that is building tailored technical solutions to modernize our government’s mission and our client’s business? Do you have a desire to change how people work? Are you interested in helping to protect our nation’s cyber interests? Join our growing team as a HPC Infrastructure & Cluster Engineer, supporting the NGAcustomer mission.

Responsibilities What you get to do every day:

You will manage the administration, health, and performance of the foundational compute environment. You will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Key Responsibilities
:

  • Cluster Administration:
    Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management:
    Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:

    AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization:
    Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management:
    Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an Infini Band GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration:
    Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat Open Shift, required for seamless customer model deployment.
  • Security and Compliance:
    Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
Qualifications

Clearance Required: Active TS clearance (with SCI Eligibility) and eligibility to obtain CI Poly

Education/

Experience:

  • Requires Bachelor's degree
  • 5+years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.

Required Skills:
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with Infini Band)
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM)
  • Hands-on experience with enterprise container orchestration platforms, specifically Open Shift or Kubernetes
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance
  • Troubleshooting Focus:
    Proven ability to diagnose and resolve complex hardware, network, and OS-level issues

What is ideal?

  • Familiarity with parallel file systems and high-throughput storage architecture.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies
  • Intelligence Community Experience preferred
What you can expect from us
  • Recognizing great…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary