Site Reliability Engineer - Elastic
Listed on 2026-08-05
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Site Reliability Engineer (SRE)
We are seeking a Site Reliability Engineer (SRE) to support mission-critical systems for U.S. Air Force programs. This role focuses on maintaining system reliability, performance, and visibility across Kubernetes-based environments.
Candidates must be eligible to obtain and maintain a DoD Secret Clearance. This position offers the opportunity to work on high-impact systems supporting national defense initiatives.
Title:
Site Reliability Engineer (SRE) – Kubernetes / Elastic (ECK Preferred)
Location Requirement:
Hybrid (3 days' on-site at any of these locations)
- Langley AFB (VA beach) or
- Hanscom AFB (Boston)
Type of Employment:
Direct Hire
- Ensure high availability, performance, and reliability of production systems
- Work hands-on with Kubernetes clusters in a mission-critical environment
- Build and maintain monitoring dashboards and visualizations (system health, performance, network bandwidth)
- Manage and update data collection/observability pipelines across clusters
- Plan and coordinate cluster maintenance, shutdowns, and data retention activities
- Create engineering diagrams and dashboards to improve system visibility and troubleshooting
- Partner with engineering, Dev Ops, and security teams to improve system stability
- Continuously improve monitoring, alerting, and automation
- Strong experience with:
- Kubernetes (required)
- Elastic Stack (Elasticsearch, Kibana, etc.)
- Experience running Elastic in a Kubernetes environment
- ECK (Elastic Cloud on Kubernetes) strongly preferred
- Elastic Cloud + Kubernetes experience is acceptable if ECK is not present
- Background in:
- Monitoring / observability / logging
- Building dashboards and system visualizations
- Experience supporting production systems (uptime, performance, reliability)
- Comfortable working in Linux-based environments
- Exposure to security tooling or SIEM platforms
- Experience with automation/scripting (Python, Bash, etc.)
- Cloud experience (AWS, Azure, or GCP)
- Experience in hybrid or on-prem environments
- Elastic certifications
- Kubernetes certifications (CKA or similar)
- Cloud certifications (AWS, Azure, etc.)
Certifications are a plus but not required.
Additional Details:
- Clearance:
Active Secret Clearance or ability to obtain and maintain one required - Location:
Onsite (VA Beach or Boston area) - Work Environment:
Enterprise-scale, mission-critical systems
What We're Looking For:
- Someone who can keep systems stable and prevent outages
- Strong Kubernetes + monitoring/observability background
- Ability to build dashboards and clearly visualize system performance
- Hands-on experience with Elastic in real environments
- Bonus if you've worked with ECK specifically
*** No 3rd Party Resumes will be accepted for this job requirement***
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).