×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer

Job in Arlington, Arlington County, Virginia, 22201, USA
Listing for: Jobtailor
Full Time position
Listed on 2026-07-20
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Support
Salary/Wage Range or Industry Benchmark: 140000 - 200000 USD Yearly USD 140000.00 200000.00 YEAR
Job Description & How to Apply Below

Responsibilities

You will own the reliability, scalability, and security of the production application and/or platform. You will do this by:

  • Implementing a World-Class Observability Platform:
    Design, implement, and manage our monitoring, logging, and alerting stack (e.g., Prometheus, Loki, Alloy, and Grafana). You won't just track metrics; you'll create the actionable insights and automated alerting that allow teams to identify and resolve issues before they impact users.
  • Defining and Upholding Reliability:
    Define, measure, and own alerting that feeds into our Service Level Indicators (SLIs) and Service Level Objectives (SLOs), increasing trust internally and externally. You will be the organization's expert on what it means for our systems to be reliable and how to measure it.
  • Leading Incident Response:
    Act as the incident responder and potentially incident commander during critical incidents who will lead blameless post-mortems / After Action Reviews (AARs) that identify true root causes and drive automated, long-term solutions to prevent recurrence.
  • Automating for Scale and Security:
    Partner with platform engineers to design, build, and manage secure, resilient Kubernetes clusters and cloud/on-prem environments using Infrastructure-as-Code (Terraform, Ansible). You will embed security and compliance controls (RMF, STIGs) directly into this automation.
  • Eliminating Toil and Scaling the Team:
    Proactively identify and eliminate operational toil by building automation. You will partner with other teams to share best practices for air-gapped environments and support their readiness for production.
Requirements
  • An active Top Secret clearance
  • 5+ years in Platform, Dev Ops, or Site Reliability Engineering with an infrastructure and operations focus.
  • Proven partner to Dev Ops/Platform and application teams; collaborates well across functions and shares context openly.
  • A deep understanding of incident response processes, with experience conducting thorough root cause analyses and driving continuous improvement.
  • Technical expertise
  • Infrastructure as Code:
    Terraform (or Cloud Formation), Ansible.
  • Containers and orchestration:
    Kubernetes design, deployment, and operations.
  • CI/CD: experience building and maintaining pipelines (Git Lab CI/CD, Jenkins, Git Hub Actions).
  • Scripting: proficiency with at least one of Python, Go, or Bash.
  • Cloud:
    Familiarity with AWS or AWS Gov Cloud.
  • Observability:
    Grafana stack, ELK stack, or Datadog.
  • Networking fundamentals: core protocols and secure configurations.
#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary