Site Reliability Engineer, Lead
Listed on 2026-07-25
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Cybersecurity
Your growth matters to us - explore our career development opportunities.
BE EMPOWERED TO SUCCEEDConnect with others in our people-first culture and enhance our collective ingenuity.
SUPPORT YOUR WELLBEINGLearn how we’ll support you as you pursue a balanced, fulfilling life.
YOUR CANDIDATE JOURNEYDiscover what to expect during your journey as a candidate with us.
As a Lead Site Reliability Engineer (SRE) on our team, you’ll be responsible for ensuring the reliability, performance, scalability, and security of critical production systems and platforms. This role leads the design and implementation of observability, automation, incident response, and operational best practices across cloud and air-gapped environments, while partnering closely with D ev O ps , infrastructure, and security teams to improve system resilience and reduce operational risk.
The Lead SRE also drives root cause analysis, capacity planning, reliability standards, and continuous improvement initiatives to support highly available , efficient, and scalable services. This is your chance to further your skills in cloud infrastructure and technologies while continuing to grow your SRE experience. You'll build and support a reliable site for the environment in order to meet the development and maintenance requirements of systems and platforms.
Work with the development and operation teams to evaluate the health, stability and reliability of systems and platforms. Design and develop technical tools to debug problems that occur in the deployment of applications, within specific platforms and systems.
Join our efforts to strengthen our security posture and safeguard national interests.
Join us. The world can’t wait.
You Have:8+ years of experience with monitoring, logging, and observability platforms, such as Prometheus, Grafana, and ELK stack
8+ years of experience with Linux systems administration and networking fundamentals within AWS
Experience with Python scripting and automation
Experience with Infrastructure as Code using Terraform and Terragrunt
Knowledge of Kubernetes administration, troubleshooting, and operations.
TS/SCI clearance with a polygraph
Bachelor’s degree and 8+ years of experience in Site Reliability Engineering, Dev Ops Engineering, or Platform Engineering, or 12+ years of experience in Site Reliability Engineering, Dev Ops Engineering, or Platform Engineering in lieu of a degree
Ability to obtain a Security+ CE, SSCP, CCNA-Security, or GSEC Certification within 6 months of start date
Nice If You Have:Experience with deploying and managing Open Telemetry .
Experience with AWS Cloud Watch, AWS EKS, and related AWS services
Experience managing Kubernetes environments through Rancher
Experience implementing SRE practices such as SLOs, SLIs, error budgets, and incident management
Experience with Jenkins, Git, Docker, Kubernetes, Nessus, JIRA, and Confluence
Knowledge of distributed systems, microservices architectures, and containerized workloads
Knowledge of NIST 800-53 and NIST-190
Master’s degree in a relevant field
Security+ CE, SSCP, CCNA-Security, or GSEC Certification
Clearance:Applicants selected will be subject to a security investigation and may need to meet eligibility requirements for access to classified information; TS/SCI clearance with polygraph is required.
CompensationAt Booz Allen, we celebrate your contributions, provide you with opportunities and choices, and support your total well-being. Our offerings include health, life, disability, financial, and retirement benefits, as well as paid leave, professional development, tuition assistance, work-life programs, and dependent care. Our recognition awards program acknowledges employees for exceptional performance and superior demonstration of our values. Full-time and part-time employees working at least 20 hours a week on a regular basis are eligible to participate in Booz Allen’s benefit programs.
Individuals that do not meet the threshold are only eligible for select offerings, not inclusive of health benefits. We encourage you to learn more about our total benefits by visiting the Resource page on our Careers site and reviewing Our Employee Benefits page.
Salary at…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).