×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Chantilly, Fairfax County, Virginia, 20153, USA
Listing for: Clearance Jobs
Full Time position
Listed on 2026-09-01
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Position: Site Reliability Engineer Jobs

Lead Site Reliability Engineer

Location:

5 days per week onsite in Chantilly, VA

Clearance Required:

TS/SCI w CI Poly required

Top Skills Senior/Lead-level Site Reliability Engineering / Production Reliability experience Strong, hands-on ELK / Elastic Stack experience:
Elasticsearch Logstash Kibana Hands-on Prometheus and/or Grafana Kubernetes

The Opportunity As a Lead Site Reliability Engineer (SRE) on our team, you'll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and implementation of observability, automation, incident response, and operational best practices across cloud and air-gapped environments, while partnering closely with Dev Ops, infrastructure, and security teams to improve system resilience and reduce operational risk.

The Lead SRE also drives root cause analysis, capacity planning, reliability standards, and continuous improvement initiatives to support highly available, efficient, and scalable services. This is your chance to further your skills in cloud infrastructure and technologies while continuing to grow your SRE experience. You'll build and support a reliable site for the environment in order to meet the development and maintenance requirements of systems and platforms.

Work with the development and operation teams to evaluate the health, stability and reliability of systems and platforms. Design and develop technical tools to debug problems that occur in the deployment of applications, within specific platforms and systems. Join our efforts to strengthen our security posture and safeguard national interests.

Qualifications 8+ years of experience with monitoring, logging, and observability platforms, such as Prometheus, Grafana, and ELK stack 8+ years of experience with Linux systems administration and networking fundamentals within AWS Experience with Python scripting and automation Experience with Infrastructure as Code using Terraform and Terragrunt Knowledge of Kubernetes administration, troubleshooting, and operations. TS/SCI clearance with a polygraph Bachelor's degree and 8+ years of experience in Site Reliability Engineering, Dev Ops Engineering, or Platform Engineering, or 12+ years of experience in Site Reliability Engineering, Dev Ops Engineering, or Platform Engineering in lieu of a degree Ability to obtain a Security+ CE, SSCP, CCNA-Security, or GSEC Certification within 6 months of start date

Nice to Have Skills Experience with deploying and managing Open Telemetry. Experience with AWS Cloud Watch, AWS EKS, and related AWS services Experience managing Kubernetes environments through Rancher Experience implementing SRE practices such as SLOs, SLIs, error budgets, and incident management Experience with Jenkins, Git, Docker, Kubernetes, Nessus, JIRA, and Confluence Knowledge of distributed systems, microservices architectures, and containerized workloads Knowledge of NIST 800-53 and NIST-190 Master's degree in a relevant field Security+ CE, SSCP, CCNA-Security, or GSEC Certification

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary