Site Reliability Engineer
Job in
Rockville, Montgomery County, Maryland, 20849, USA
Listed on 2026-07-31
Listing for:
Jobtailor
Full Time
position Listed on 2026-07-31
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, AWS, Systems Engineer
Job Description & How to Apply Below
Responsibilities
- Join the team supporting the Centers for Medicare & Medicaid Services (CMS) as it merges and modernizes its enterprise knowledge and data systems into a single, AI-driven platform, reducing manual effort, improving data accuracy, and enhancing transparency for stakeholders.
- Keep the systems up and the users happy. Operate and tune AWS environments to meet infrastructure and application availability SLAs, even during transition and change.
- Build observability that actually informs. Implement continuous monitoring, alerting, and dashboards using tools like AWS Cloud Watch, New Relic, and Splunk, and establish performance baselines so you can spot degradation before users do.
- Automate the toil. Write infrastructure-as-code (Terraform, Ansible) and support CI/CD pipelines (Jenkins) and containerized workloads (Docker) for repeatable, reliable deployments.
- Define and track the numbers that matter. Set and monitor SLIs and SLOs, and produce performance, load/stress, and bottleneck reports that drive smarter decisions.
- Optimize for performance, security, and cost. Use tools like AWS Trusted Advisor to find and act on improvement opportunities.
- Support security and compliance modernization. Partner with the Security & Compliance SME to review vulnerability and security scans, feed continuous monitoring, and help advance the move toward a Continuous ATO (cATO) within a FISMA Moderate boundary (RMF, ARS, IS2P2).
- Strengthen resilience. Help design and maintain disaster recovery and COOP continuity so the systems hold up against outages, incidents, and the unexpected.
- Own incidents end to end. Drive response, run blameless post-mortems, and implement the preventative fixes that keep the same thing from happening twice.
- A bachelor’s degree in computer science, engineering, or a related field (or equivalent hands‑on experience).
- 3–5 years of experience in site reliability, systems, or cloud engineering, with meaningful time spent in AWS environments.
- Solid working knowledge of core AWS services, architecture, and best practices.
- Hands‑on experience with infrastructure-as-code tools (Terraform, Ansible, or Cloud Formation).
- A good understanding of CI/CD pipelines and automation tools (Jenkins, Git Lab CI, or similar).
- Comfort scripting and automating in Python.
- Familiarity with monitoring and observability tooling (Cloud Watch, New Relic, Splunk, or comparable).
- Strong problem‑solving instincts and the composure to work calmly under pressure.
- Clear communication skills, with the ability to make complex technical concepts understandable.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×