Federal Platform Engineer
Listed on 2026-07-19
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Platform Engineer
The Federal Platform Engineer is a crucial member of the team who is responsible for the Agentic AI platform's scalability and reliability. Serving as the primary architect and operator of the infrastructure required to run complex, agentic applications, you will own the end-to-end lifecycle of the platform—from initial infrastructure-as-code provisioning to high-scale, day-two operations. In this role, you bridge the gap between cloud-native engineering and the unique demands of AI/LLM workloads, ensuring that agentic systems are backed by resilient, secure, and performant environments.
You are deeply passionate about automation, Kubernetes orchestration, and the operational rigor required to support real-time, stateful agent execution. Rather than just deploying "black box" infrastructure, you thrive on solving the complexities of high-scale systems: managing GPU orchestration, optimizing vector database performance, and implementing robust observability pipelines that provide visibility into the "thought process" of live agents.
Scalable Foundations:
Design and implement the path from local development to deployed federal cloud clusters (AWS Gov Cloud/Azure Government) using Kubernetes EKS/AKS, Infrastructure as Code, Terraform/Open Tofu/CDK, and Git Ops controllers Flux/Argo CD.
Optimized Compute:
Configure and manage specialized compute resources, including GPU-accelerated node groups and serverless scaling policies, to support inference and training workloads.
Network & Security Hardening:
Implement Zero Trust architecture principles, ensuring secure communication between microservices, agentic tools, and external APIs through service meshes (e.g., Istio) and egress controls.
Packaging & Promotion:
Author and maintain Helm charts for application delivery, managing the promotion path from local-to-AWS and ensuring consistency across environments.
Minimum Qualifications
Current, active TS/SCI security clearance with polygraph.
Current, active Security+ or equivalent certificate for privileged user access.
5 years of professional infrastructure, Dev Ops or platform engineering experience, with a focus on cloud-native infrastructure.
3 years of experience managing production Kubernetes environments.
3 years of experience with Flux (or equivalent Git Ops controller like Argo CD) and authoring/maintaining Helm charts.
Proven experience with Infrastructure as Code (Terraform) and CI/CD tools (Git Lab CI, Git Hub Actions).
Practical experience operating systems in U.S. Federal or DoD environments.
Everforth Apex is a world-class IT services company that serves thousands of clients across the globe. When you join Everforth Apex, you become part of a team that values innovation, collaboration, and continuous learning. We offer quality career resources, training, certifications, development opportunities, and a comprehensive benefits package. Our commitment to excellence is reflected in many awards, including Clearly Rateds Best of Staffing® in Talent Satisfaction in the United States and Great Place to Work® in the United Kingdom and Mexico.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).