×
Register Here to Apply for Jobs or Post Jobs. X

AWS DevOps​/Agentic SRE Engineer; TS​/SCI Clearance

Job in Chantilly, Fairfax County, Virginia, 22021, USA
Listing for: Strategic Business Systems
Full Time position
Listed on 2026-09-09
Job specializations:
  • IT/Tech
    AWS, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 180000 - 270000 USD Yearly USD 180000.00 270000.00 YEAR
Job Description & How to Apply Below
Position: AWS DevOps / Agentic SRE Engineer (Active TS/SCI Clearance)

AWS Dev Ops / Agentic SRE Engineer TS/SCI

Clearance: Active TS/SCI
Certification: Active Security+ or equivalent required
Experience: 3+ years of SRE, Dev Ops, Platform Engineering, or Infrastructure Engineering experience

Position Overview

We are seeking an SRE / Dev Ops & Release Engineer to own the deployment, reliability, observability, and operation of secure AWS environments supporting an advanced Agentic AI platform.

This is a hands-on engineering position combining Dev Ops, Site Reliability Engineering, platform engineering, and release automation. The engineer will own the path from local development through production deployment into AWS Kubernetes environments, as well as the operational health and reliability of those environments.

The platform is currently operating within IATT environments and progressing toward scale and ATO. At the current team size, build/release and production operations are intentionally combined the engineer responsible for deploying the platform also has responsibility for ensuring that it operates reliably.

The environment includes AWS EKS, CDK, Kubernetes, Helm, Flux Git Ops, container registries, Envoy Gateway, distributed tracing, and modern observability technologies.

Key Responsibilities
  • Design, build, maintain, and operate highly available AWS infrastructure supporting an Agentic AI platform.
  • Own the software delivery lifecycle from local development through build, test, packaging, promotion, and deployment into AWS environments.
  • Deploy and operate Kubernetes workloads in production AWS EKS environments.
  • Author, maintain, and troubleshoot Helm charts supporting platform applications and services.
  • Build and maintain Git Ops-based continuous delivery utilizing Flux, Helm, and container registries.
  • Develop and maintain Infrastructure as Code using AWS CDK and Type Script.
  • Build and manage AWS infrastructure utilizing EKS, RDS, S3, IAM/IRSA, ECR, and related AWS services.
  • Build and maintain container-image and Helm-chart promotion processes.
  • Configure and support Envoy Gateway routing, certificates, and TLS.
  • Implement and operate observability and distributed-tracing capabilities utilizing Open Telemetry, Grafana, Tempo, or equivalent technologies.
  • Establish cross-service tracing to provide end-to-end visibility into distributed application and agent workflows.
  • Monitor cluster, application, and service health and proactively identify reliability issues.
  • Diagnose failed or stuck Helm and Flux reconciliations, configuration drift, pod failures, and deployment issues.
  • Lead production incident investigation and root-cause analysis using Kubernetes state, logs, metrics, and distributed traces.
  • Develop preflight validation, diagnostic, QA, and operational tooling.
  • Automate infrastructure and operational processes using Python and other scripting technologies.
  • Support credential, certificate, and secrets rotation.
  • Partner with software engineers and AI/ML teams to deploy and operate agentic applications, model-serving infrastructure, and supporting services.
  • Support platform security, compliance, IATT, and ATO activities.
  • Use AI-assisted development tools to accelerate engineering while independently validating generated code and configuration before deployment.
  • Drive an engineering approach based on measurable evidence, automated testing, telemetry, and verification.
Required Qualifications
  • Current and active TS/SCI security clearance
  • Current Security+ certification or equivalent certification supporting privileged-user access.
  • 3+ years of professional experience in Site Reliability Engineering, Dev Ops, Platform Engineering, Cloud Engineering, or Infrastructure Engineering.
  • Hands-on experience operating Kubernetes in production environments.
  • Experience authoring and maintaining Helm charts.
  • Experience implementing Git Ops-based continuous delivery using Flux, Argo CD, or equivalent technologies.
  • Hands-on AWS experience with services such as EKS, RDS, S3, IAM/IRSA, and ECR.
  • Experience developing and maintaining Infrastructure as Code.
  • Experience working with Type Script/AWS CDK or demonstrated ability to work within a Type Script-based IaC environment.
  • Experience implementing and utilizing production observability and distributed-tracing technologies such as Open Telemetry, Grafana, and Tempo.
  • Demonstrated experience diagnosing infrastructure and application failures using telemetry, logs, metrics, and traces.
  • Experience leading incident investigation, root-cause analysis, remediation, and validation.
  • Proficiency with Python or another scripting language for infrastructure automation and operational tooling.
  • Strong understanding of CI/CD, containers, networking, security, and modern cloud architecture.
Preferred Qualifications
  • TS/SCI with Poly
  • Experience implementing registry-based Git Ops architectures utilizing ECR or similar container registries.
  • Experience troubleshooting Flux and Helm reconciliation issues, including Helm Release failures and configuration drift.
  • Experience deploying and operating ML/LLM-serving…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary