×
Register Here to Apply for Jobs or Post Jobs. X

SME Platform Engineer

Job in Arlington, Arlington County, Virginia, 22201, USA
Listing for: General Dynamics Information Technology
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Cybersecurity, Systems Engineer
Salary/Wage Range or Industry Benchmark: 150000 - 190000 USD Yearly USD 150000.00 190000.00 YEAR
Job Description & How to Apply Below

Type of

Requisition :
Regular Clearance Level Must Currently Possess:
Top Secret/SCI Clearance Level Must Be Able to Obtain:
Top Secret SCI + Polygraph Public Trust/Other

Required:

None Job Family: IT Infrastructure and Operations

Job Qualifications:

Skills:

CI/CD, Cloud Infrastructure, Cluster Administration, Kubernetes, Linux Server Administration

Certifications:

None

Experience:

15 + years of related experience US Citizenship

Required:

Yes

Job Description:

YOUR IMPACT

Own your opportunity to work with the largest government agency in the nation. Make an impact by advancing the Department of War’s mission to keep our country safe and secure.

OUR COMPANY Iron EagleX (IEX), a wholly owned subsidiary of General Dynamics Information Technology (GDIT), delivers agile IT and Intelligence solutions. Combining small-team flexibility with global scale, IEX leverages emerging technologies to provide innovative, user-focused solutions that empower organizations and end users to operate smarter, faster, and more securely in dynamic environments.

JOB DESCRIPTION

Iron EagleX is seeking a SME Platform Engineer to support our Engineering team in Crystal City, VA. This role will lead the design, implementation, and management of our secure, on-premises cloud infrastructure. In this role, you will be the driving force behind our advanced computing environments, ensuring the seamless orchestration of containerized applications and large-scale data science platforms. You will work at the intersection of infrastructure, security, and machine learning, managing robust compute clusters and providing foundational support for AI model training and deployment.

The ideal candidate has deep expertise in Kubernetes ecosystem tools, Git Ops methodologies, and strict compliance standards.

MEANINGFUL WORK AND PERSONAL IMPACT

As a SME Platform Engineer, your work will directly empower our data science and engineering teams to push the boundaries of machine learning and data analytics. By building and maintaining resilient GPU and Ray clusters, you will accelerate the fine-tuning and deployment of advanced models. Your commitment to security and compliance will ensure our critical systems remain protected against vulnerabilities, providing a safe, compliant, and highly performant foundation for the organization’s most impactful technical initiatives.

You will not just be managing infrastructure; you will be enabling innovation.

JOB DUTIES (INCLUDE BUT ARE NOT LIMITED TO)
  • Infrastructure & Orchestration:
    Architect, deploy, and manage on-premises cloud infrastructure using RKE2 and maintain storage solutions like Longhorn and Object storage.
  • Platform Enablement:
    Host and maintain robust data science environments, including software such as POSIT Workbench/Connect and Hive Metastore.
  • AI/ML Infrastructure:
    Manage and scale robust GPU clusters, Ray Clusters for fine-tuning machine learning models, and VLLM Routers for efficient model inference.
  • CI/CD & Automation:
    Build, maintain, and optimize CI/CD pipelines using Git, Helm charts, and ArgoCD for reliable software delivery.
  • Security & Compliance:
    Ensure continuous FIPS compliance across the environment. Actively manage and mitigate critical and high-level vulnerabilities.
  • Identity & Access:
    Implement and maintain robust authentication and authorization mechanisms using Keycloak and Open Policy Agent (OPA).
  • System Administration:
    Pull and manage container images from secure registries such as Harbor, Docker Hub, or Container yard. Manage all core capabilities and troubleshoot issues effectively via the command-line console.
REQUIRED SKILLS
  • Demonstrated experience designing, deploying, administering, and troubleshooting production Kubernetes environments; hands-on experience with RKE2 or similar.
  • Strong Linux systems administration skills, including the ability to manage, diagnose, and troubleshoot infrastructure and platform services through the command line (CLI).
  • Experience implementing authentication and authorization solutions using technologies such as Keycloak, Open Policy Agent (OPA), OIDC, RBAC, or comparable identity and access management frameworks.
  • Hands-on experience with Git-based development and deployment workflows, including Helm charts, CI/CD pipelines, and Git Ops practices.
  • Experience with Argo CD or similar tools for declarative, Git Ops-based continuous delivery.
  • Experience managing Kubernetes storage solutions, including distributed block storage and object storage; experience with Longhorn or comparable…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary