AI Platform Engineer - Chantilly, VA - TS/SCI CI Polygraph
Listed on 2026-07-15
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Infrastructure
Supporting the Federal Government, contractors and subcontractors by delivering the industry’s best cleared talent and staffing resources.
Added - 07/09/26 AI Platform Engineer - Chantilly, VA - TS/SCI CI Polygraph Required Engineering Chantilly, Virginia | Contract To Hire
AI Platform Engineer (Long-Term Contract) – Chantilly, VACandidate must have an active TS/SCI CI Polygraph clearance.
Job SummaryAs a key member of the infrastructure support team, the Mid-Level AI Platform Engineer will be responsible for designing, implementing, and maintaining highly available, scalable infrastructure platforms. This position plays a critical role in supporting engineers responsible for leveraging next‑generation infrastructure to drive innovation in the Artificial Intelligence (AI) field. You will act as the bridge between platform infrastructure and software engineering teams, delivering standard “paved roads” and internal developer platforms for secure, high‑tempo AI workloads.
Key Responsibilities- Cluster Management: design, deploy, secure, maintain, and upgrade highly available, scalable Kubernetes clusters (EKS, AKS, GKE, Open Shift, or self‑managed environments).
- Pipeline Design & Implementation: design, implement, and maintain robust, automated CI/CD pipelines for deploying containerized applications to Kubernetes, adhering to modern Git Ops workflows.
- Infrastructure as Code (IaC): provision and manage distributed cloud infrastructure declaratively using standard IaC frameworks and toolsets.
- Automation & Custom Tooling: write robust custom scripts and tooling to automate routine operational tasks, eliminate engineering bottlenecks, and integrate diverse infrastructure components.
- Performance Tuning & Optimization: diagnose and resolve complex performance issues across the entire Kubernetes stack, optimize multi-tenant resource utilization, and perform regular cluster tuning.
- Monitoring & Observability: implement and maintain comprehensive monitoring, logging, and alerting solutions to capture real‑time platform metrics and distribute tracing data.
- Systems Hardening & Compliance: secure distributed platform environments by enforcing strict access boundaries, managing platform secrets, and triaging vulnerabilities in line with federal security frameworks.
- Field Support & Travel: travel up to 20% of the time, as required, to perform on‑site installations, maintenance, and troubleshooting activities at customer sites or data centers.
- Containerization & Orchestration: proficient with Docker and advanced Kubernetes primitives including Pods, Deployments, Stateful Sets, Services, Ingress, Config Maps, Secrets, Persistent Volumes, and Name spaces.
- Git Ops & Continuous Delivery: practical experience with Git Ops methodologies and automated continuous delivery deployment tools like Argo CD.
- CI/CD Tooling: hands‑on experience with popular automation platforms such as Jenkins, Git Lab CI/CD, Git Hub Actions, Tekton, or Argo Workflows.
- Application Packaging: experience building and managing configurations using Helm charts and Kustomize for standardized Kubernetes application packaging.
- Core Infrastructure Networking: deep understanding of cluster networking concepts (CNI plugins like Calico or Cilium), network policies, service meshes (Istio, Linkerd), alongside fundamental protocols (TCP/IP, DNS, HTTP, Load Balancing).
- Platform Security: expertise implementing Role‑Based Access Control (RBAC), Pod Security Policies/Admission Controllers, vulnerability scanning, and secure secret management (Vault, AWS Secrets Manager).
- Linux/Unix Systems: strong background in Linux system administration, including networking configurations, file systems, process management, and rapid troubleshooting.
- Problem‑Solving: excellent analytical skills to isolate, diagnose, and resolve complex technical flaws across complex distributed software stacks.
- Communication &
Collaboration:
strong verbal and written communication skills to collaborate effectively with development teams, operations, security, and external program stakeholders. - Proactive Ownership: a forward‑thinking approach to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).