Senior Site Reliability Engineer with Security Clearance
Job in
Fairfax, Fairfax County, Virginia, 22031, USA
Listed on 2026-08-11
Listing for:
ECS
Full Time
position Listed on 2026-08-11
Job specializations:
-
IT/Tech
Cybersecurity, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
ECS is responsible for designing, building, deploying, operating, and maintaining a complete 'Data Services' solution which includes the collection, normalization, visualization, and sharing of cyber data from more than 100 Federal agencies. The CDM Data Services product is an integrated suite of multiple Commercial Off the Shelf (COTS) products, software configuration packages, and custom code which work together to operate as an integrated solution tailored to meet Department of Homeland Security (DHS) requirements. 
We are seeking professionals who thrive in a dynamic, fast-paced, and highly collaborative environment where problem-solving, critical thinking, and a holistic approach to serving the mission are key. Our program operates within the Scaled Agile Framework (SAFe). An aptitude and enthusiasm for continuous learning, improvement, and cyber security is a must! Role & Responsibilities : ECS is seeking a talented Senior Site R eliability Engineer ( SRE ) to play a key role in defining, implementing, and growing our SRE practice to ensure the reliability, availability, and performance of our critical production environments.
The Senior SRE will contribute to a culture of continuous improvement, identifying areas for enhancement, and driving initiatives to improve system reliability, scalability, and efficiency . The successful candidate will have demonstrated hands-on experience design ing , implement ing , and maintain ing solutions to ensure that systems, including infrastructure and applications , are resilient, highly available , and performant .
The Senior SRE will also play a critical role in defining and measuring the Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for our solution. The Senior SRE will be responsible for set ting up comprehensive logging, monitoring , and alerting s olutions using the Elastic s tack and other tools as necessary to ensure the continuous performance of services .
Additionally, they will respond to incidents, perform root cause analys e s, and implement solutions to prevent re o c c urrence s . The Senior SRE will work in close collaboration with other SRE team members, d evelopers, t esters, i nfrastructure engineers, Dev Ops engineers, and other stakeholders to integrate reliability and observability into the software development lifecycle.
Salary Range: $118,000 - $177,000 Required Skills
* Must be a US citizen with the ability to obtain Public Trust Suitability .
* 6 + years of experience as a Site Reliability Engineer (SRE) or equivalent
* 6 + years of demonstrated experience designing, implementing , and maintaining observability solutions to include logging, monitoring, and alerting
* 6 + years of hands-on experience with SRE tools (e.g., Elastic, Prometheus, Grafana, Splunk , etc.)
* 3 + years defining and measuring SLOs and SLIs
* 3 + years of relevant experience using cloud platforms (AWS Gov Cloud preferred )
* 3 + years of hands-on programming or scripting (e.g., Python, Bash, etc.)
* Strong knowledge of microservices, containerization, and orchestration tools (Docker, Kubernetes)
* Proven ability to collaborate with cross-functional teams (development, testing, and product) to integrate reliability and observability into the software development lifecycle
* Strong problem-solving and analytical skills
* Proactive, detail-oriented approach to identifying inefficiencies and implementing improvements .
* Proficient in developing Synthetic monitoring scripts using typescript. Desired Skills
* Bachelor's degree in Computer Science , Engineering, or a related field (or 4 additional years of related experience)
* Experience working in an Agile/ SAFe environment using ALM tools (Jira, Confluence, or similar)
* Strong understanding of CI/CD principles and platforms (Jenkins, CircleCI , Git Lab, Git Hub Actions, Argo, Travis CI, etc.)
* Expertise in configuration management tools (Ansible, Puppet, Chef)
* Experience with infrastructure as code (Terraform, Cloud Formation)
* In-depth understanding of networking, security, and system administration of Linux operating systems
* Knowledge of version control platforms and branching strategies
* Knowledge of disaster recovery planning, backup strategies, and data replication
* Experience supporting large Federal programs ($200M+) #EverforthECS1 ECS Federal LLC is an equal opportunity employer and does not discriminate or allow discrimination on the basis any…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×