Site Reliability Engineer - Disaster Recovery & Business Continuity
Job in
Brockton, Plymouth County, Massachusetts, 02411, USA
Listed on 2026-07-19
Listing for:
Charles River Associates
Full Time
position Listed on 2026-07-19
Job specializations:
-
IT/Tech
SRE/Site Reliability, Disaster Recovery IT, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Position Overview
The Site Reliability Engineer (SRE) helps ensure CRA’s critical business services are reliable, scalable, and performant across on-premises and cloud environments. This role blends software engineering and operations practices to reduce manual toil through automation, improve service observability, and strengthen incident response. The SRE partners closely with infrastructure, security, application, and service delivery teams to define measurable reliability targets (SLIs/SLOs), implement resilient architectures, and drive continuous improvement through blameless post‑incident learning.
Key Responsibilities- Hands‑on system engineering experience with core enterprise infrastructure platforms and services, including Windows Server, VMware vSphere, VMware Site Recovery Manager (SRM), SAN technologies, and the Rubrik ecosystem, with the ability to understand dependencies, recovery workflows, and failure modes across on‑premises and cloud environments
- Service ownership & reliability targets: partner with service owners to define and maintain SLIs and SLOs for availability, latency, and performance; track error budgets and reliability risk.
- Observability: implement and continuously improve monitoring, logging, alerting, and dashboards to provide actionable, symptom‑based signals and reduce mean time to detect/respond (MTTD/MTTR).
- Blameless post‑mortems & continuous improvement: facilitate post‑incident reviews, identify root causes and contributing factors, and drive remediation items to completion; standardize learnings into runbooks and operational practices.
- DR testing program build‑out: design and launch a scalable DR testing program (scope, test types, cadence, success criteria, and evidence capture) in partnership with application, infrastructure, and security teams; maintain runbooks and lead regular tabletop and technical recovery exercises to validate RTO/RPO assumptions and improve recoverability.
- DR readiness: contribute to reliability architecture and disaster recovery readiness for key services, including dependency mapping, recovery testing inputs, and validation of recovery procedures.
- Cross‑functional collaboration: work day‑to‑day with infrastructure, network, cloud, security, and application teams to improve operational excellence, reliability culture, and shared ownership of production outcomes.
- Experience operating and improving reliability of production services (on‑prem and/or cloud), including incident response, operational readiness, and service ownership
- Working knowledge of SRE concepts and practices such as SLIs/SLOs, error budgets, monitoring/alerting strategy, and blameless postmortems
- Experience with observability tooling and practices (logs, metrics, tracing, dashboards) and using data to drive reliability and performance improvements
- Experience with disaster recovery orchestration and recovery testing using VMware Site Recovery Manager (SRM) and Azure Site Recovery (ASR) (or similar public cloud DR services)
- Proven experience building and operating a DR testing program, including dependency mapping, test planning, coordination across stakeholders, execution of tabletop and technical failover tests, documentation of results, and tracking remediation actions to closure
- Strong cross‑functional communication and teamwork skills; comfortable partnering with engineering, security, and operations teams to drive shared outcomes
- Ability to document and standardize operational procedures (runbooks), participate in on‑call rotations, and manage multiple priorities in a fast‑moving environment
- CRA’s robust skills development programs, including a commitment to offering 100 hours of training annually through formal and informal programs, encourage you to thrive as an individual and team member. Beginning with research and analysis skill building, training continues with technical training, presentation skills, internal seminars, and career mentoring and performance coaching from an assigned senior colleague. Additional leadership and collaboration opportunities exist through internal firm development activities.
- We offer…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×