Site Reliability Engineer
Listed on 2026-07-09
-
IT/Tech
AWS, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Skyward is a people‑centered business dedicated to improving human advancement through information technology. We serve our clients with a value‑driven, collaborative culture that emphasizes integrity, compassion, inquisitiveness, and authenticity. We are seeking an experienced Site Reliability Engineer to join our team.
What You'll Do- Support the Centers for Medicare & Medicaid Services (CMS) as it merges and modernizes its enterprise knowledge and data systems into a single, AI‑driven platform. Reduce manual effort, improve data accuracy, and enhance transparency for stakeholders.
- Operate and tune AWS environments to meet infrastructure and application availability SLAs, even during transition and change.
- Build observability that informs; implement continuous monitoring, alerting, and dashboards using AWS Cloud Watch, New Relic, and Splunk; establish performance baselines.
- Automate toil; write infrastructure‑as‑code (Terraform, Ansible) and support CI/CD pipelines (Jenkins) and containerized workloads (Docker).
- Define and track SLI and SLO metrics; produce performance, load/stress, and bottleneck reports.
- Optimize for performance, security, and cost; use tools such as AWS Trusted Advisor.
- Partner with Security & Compliance SME to review vulnerability scans, feed continuous monitoring, and advance the move toward continuous ATO within a FISMA Moderate boundary.
- Design and maintain disaster recovery and COOP continuity plans.
- Own incidents end to end; drive response, run blameless post‑mortems, and implement preventative fixes.
- Bachelor’s degree in computer science, engineering, or a related field (or equivalent experience).
- 3–5 years of site reliability, systems, or cloud engineering experience, with significant time in AWS.
- Solid knowledge of core AWS services, architecture, and best practices.
- Hands‑on experience with infrastructure‑as‑code tools (Terraform, Ansible, or Cloud Formation).
- Experience with CI/CD pipelines and automation tools (Jenkins, Git Lab CI, etc.).
- Comfort scripting in Python.
- Familiarity with monitoring and observability tooling (Cloud Watch, New Relic, Splunk, etc.).
- Strong problem‑solving instincts and composure under pressure.
- Clear communication skills; ability to explain complex technical concepts simply.
- Experience working with CMS.
- Experience in AI, NLP, or LLM‑driven environments.
- AWS certifications and the real‑world experience that comes with them.
- Medical, dental, vision insurance (fully paid for employees)
- 15 days of paid leave
- 7 days of sick leave
- 2 days bereavement leave
- 11 paid Federal holidays
- Up to 40 hours for jury duty
- 401(k) with 4% employer contribution (no vesting period)
- Up to 4 weeks of paid paternity and maternity leave
- Company‑provided laptop
- $5,000 per year for professional development
- $600 per year for technical supplies and equipment
- $2,000 referral bonus
- Life and disability insurance
- HSA and FSA
- Legal Shield and d Voluntary Benefits
- Opportunity to work in a collaborative, motivated team focused on modernizing government services with cutting‑edge technology.
$112,000 - $150,000 a year. We believe great work deserves great pay. Compensation is competitive, transparent, and fair, as required by Maryland law. We support flexible working hours and remote opportunities for a healthy work‑life balance.
Offers of employment are contingent upon acceptable background investigation results. Applicants must obtain and maintain a Public Trust security clearance due to the nature of our work as a government contractor. We may use artificial intelligence tools in the hiring process, but final hiring decisions are made by human recruiters.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).