×
Register Here to Apply for Jobs or Post Jobs. X

HPC Infrastructure & Cluster Engineer

Job in Springfield, Fairfax County, Virginia, 22161, USA
Listing for: Inflowfed
Full Time position
Listed on 2026-09-06
Job specializations:
  • IT/Tech
    Systems Engineer, IT Infrastructure, Systems Administrator, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 120000 - 150000 USD Yearly USD 120000.00 150000.00 YEAR
Job Description & How to Apply Below

About INflow Federal - founded in 2013, INflow Federal is a mission-driven small business delivering cutting-edge solutions to the Department of War (DoW) and Joint Force operations across 20+ states. Our strength comes from our people - especially the Veterans who make up over 50% of our workforce. Through our Veteran Outreach Program and employee-first culture, we invest deeply in professional growth, well-being, and innovation.

Known for our agility, transparency, and integrity, INflow combines real-world experience with emerging technologies like AI/ML to help our customers lead in a rapidly evolving defense landscape. We empower both our employees and mission partners to stay ahead - driving smarter, faster, and more secure outcomes.

Job Overview:

We are seeking an Infrastructure & Cluster Engineer to manage the administration, health, and performance of the foundational compute environmen. In this role, you will be responsible for the end-to-end administration of a dedicated customer compute cluster. Your primary mission is to ensure a highly available, secure, and optimized hardware foundation. By maintaining a robust infrastructure, you will directly contribute to the critical technology integration and performance engineering efforts, ensuring a highly reliable platform for integrating and executing complex customer workloads.

Here, your work is more than a job- it's a journey in innovation. With opportunities to work on high-impact projects, access to the latest technologies, and a culture that thrives on creativity and collaboration, INflow Federal is where your expertise can truly make a difference.

  • Cluster Administration:

    Manage the day-to-day operations of the customer compute cluster, including Linux operating system administration, hardware monitoring, patching, and system upgrades.
  • Resource and Job Management:

    Configure, maintain, and optimize workload management and orchestration platforms, utilizing the Run:

    AI job scheduler to ensure efficient distribution of intensive AI/ML workloads across the cluster.
  • Infrastructure Optimization:

    Tune cluster performance at the hardware, operating system, and network levels to maximize compute efficiency and data throughput for customer workloads.
  • Storage and Network Management:

    Administer storage solutions and high-speed networking fabrics. Support the transition to and ongoing management of an Infini Band GPU-to-GPU network infrastructure to minimize latency for distributed operations.
  • Environment Configuration:

    Partner with technology integration teams to provision specific environments, dependencies, and container platforms, specifically leveraging Red Hat Open Shift, required for seamless customer model deployment.
  • Security and Compliance:

    Ensure all infrastructure components remain compliant with federal security standards, implementing strict access controls and maintaining system accreditations.
  • Experience:

    5+ years of experience in Linux systems administration and infrastructure management with a specific focus on high-performance computing environments.
  • Technical

    Skills:
  • Expertise in managing bare-metal servers, enterprise storage arrays, and advanced network configurations (Experience with Infini Band).
  • Strong proficiency with workload managers, job schedulers, and AI orchestration tools (e.g., Run:AI, SLURM).
  • Hands-on experience with enterprise container orchestration platforms, specifically Open Shift or Kubernetes.
  • Experience writing automation and configuration scripts (e.g., Bash, Python) to streamline cluster maintenance.
  • Troubleshooting Focus:

    Proven ability to diagnose and resolve complex hardware, network, and OS-level issues.
  • Familiarity with parallel file systems and high-throughput storage architectures.
  • Prior experience engineering or managing high-speed GPU-to-GPU communication topologies.
Certification Requirements

We built 8140.study - our certification training platform - to help. This helps you study for certifications such as the CompTIA Security+.

Other Notes
  • Some travel may be required:
    Must have valid driver’s license and transportation. This is subject to change at the direction of the customer.
  • If accommodation is needed with your…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary