Site Reliability Engineer
Job in
Baltimore, Anne Arundel County, Maryland, 21283, USA
Listed on 2026-10-01
Listing for:
Leidos
Full Time
position Listed on 2026-10-01
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, AWS, SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below
Primary Responsibilities The SRE/Senior Cloud Engineer shall design, build, and operate the highly available cloud infrastructure supporting the CRM modernization effort.
Responsibilities shall include, but are not limited to:
Design, build, and operate highly available AWS infrastructure within the CMS AWS Enclave (FedRAMP Moderate), applying AWS Well-Architected Framework best practices.
Architect secure, scalable multi-account VPC interconnectivity between the CMS AWS Enclave and Pega Cloud for Government (PCFG) in AWS Gov Cloud US-West, including Private Link, API Gateway, and Direct Connect.
Support container and serverless architectures (e.g., AWS Lambda, Glue) for data integration, batch processing, and API layers supporting the modernized CRM.Apply Site Reliability Engineering (SRE) practices, including defining and tracking Service Level Objectives (SLOs), error budgets, and reliability metrics aligned to contract SLAs (e.g., 99.9% availability).Build and maintain observability across the AWS and Pega environments, including centralized logging, monitoring, and alerting, to enable proactive detection of performance and availability issues.
Lead incident response and root-cause analysis for production issues, and long-term reliability improvements.
Automate infrastructure provisioning, configuration, and environment build-out using Infrastructure as Code (e.g., Terraform, Cloud Formation, Ansible).Design and test Disaster Recovery capability for cloud-based workloads, including backup, failover, and Multi-AZ/Multi-Region resilience.
Support performance testing and capacity planning to validate the platform's ability to scale to 20,000 concurrent CSR sessions and peak Open Enrollment Period (OEP) volumes.
Support continuous security monitoring, vulnerability remediation, and Zero Trust alignment across the AWS and Pega environments.
Partner with the Dev Ops Lead/Configuration Manager to build and maintain CI/CD pipelines, ensuring automated testing, security scanning, and deployment across all SDLC environments.
Support cloud connectivity and data movement for the AWS-based data migration pipeline (e.g., AWS Glue, S3, RDS/Aurora PostgreSQL) between legacy Siebel and the modernized CRM.Coordinate with the CMS Hybrid Cloud Team on cloud environment provisioning, patching, and lifecycle management activities.
Manage release coordination and change windows in support of OEP blackout periods and other critical operational periods, minimizing risk of service disruption.
Collaborate with the Release Train Engineer, Solution Architect, and Agile delivery teams to align infrastructure readiness with sprint and PI planning.
Support integration of Genesys Cloud CX infrastructure and telephony/chat channels with the modernized CRM environment.
Continuously identify opportunities to reduce operational toil through automation of repetitive tasks, log analysis, and routine operational activities.
Document cloud architecture, operational runbooks, and disaster recovery procedures to support the Transition-Out Plan and audit readiness.
Provide technical mentoring and knowledge-sharing to other engineers on cloud architecture, automation, and reliability engineering best practices.
Communicate technical status, risks, and dependencies to CMS leadership, and Leidos management.
Support requirements traceability and technical documentation related to infrastructure and integration architecture.
Actively participate in planning sessions, requirements gathering activities, design sessions, Agile sessions, and other events supporting the CRM modernization effort.
Required Qualifications:
Bachelor’s degree and a minimum of 6-8 years of relevant experience in cloud engineering, site reliability engineering, or infrastructure, or an equivalent combination of education and experience
Experience building highly available AWS infrastructure based on industry best practices and the AWS Well-Architected Framework
Experience with Infrastructure as Code, automation, and configuration management of cloud-based resources
Experience designing Disaster Recovery for cloud-based workloads,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×