More jobs:
Senior Site Reliability Engineer
Job in
Arlington, Arlington County, Virginia, 22201, USA
Listed on 2026-07-20
Listing for:
Jobtailor
Full Time
position Listed on 2026-07-20
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Support
Job Description & How to Apply Below
Responsibilities
You will own the reliability, scalability, and security of the production application and/or platform. You will do this by:
- Implementing a World-Class Observability Platform:
Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics; you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users. - Defining and Upholding Reliability:
Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally. You will be the organization's expert on what it means for our systems to be reliable and how to measure it. - Leading Incident Response:
Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence. - Automating for Scale and Security:
Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible). You will embed security and compliance controls (RMF, STIGs) directly into this automation. - Eliminating Toil and Scaling the Team:
Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production.
- An active Top Secret clearance
- 5+ years in Platform, Dev Ops, or Site Reliability Engineering with an infrastructure and operations focus.
- Proven partner to Dev Ops/Platform and application teams; collaborates well across functions and shares context openly.
- A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement.
- Technical expertise
- Infrastructure as Code:
Terraform (or Cloud Formation), Ansible. - Containers and orchestration:
Kubernetes design, deployment, and operations. - CI/CD: experience building and maintaining pipelines (Git Lab CI/CD, Jenkins, Git Hub Actions).
- Scripting: proficiency with at least one of Python, Go, or Bash.
- Cloud:
Familiarity with AWS or AWS Gov Cloud. - Observability:
Grafana stack, ELK stack, or Datadog. - Networking fundamentals: core protocols and secure configurations.
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×