Principal Systems Engineer, Enterprise Backup & Disaster Recovery Solutions
Listed on 2026-08-31
-
IT/Tech
Disaster Recovery IT, Systems Engineer, Cloud Computing: Infrastructure & Operations
Principal Systems Engineering Role
This is a hands-on, Principal Systems Engineering role responsible for leading the design, implementation, operation and continuous improvement of enterprise data protection and disaster recovery (DR) solutions across cloud, on-prem, hybrid, and scientific systems infrastructure. The Principal Systems Engineer scales backup and recovery coverage across enterprise and lab systems, establishes and operates SOPs to regularly test backups and validate recoverability. The individual in this role defines and implements DR solutions for core infrastructure systems and services.
This includes building the foundational DR capability from the ground up, planning and executing DR exercises on a recurring cadence, and maintaining backup, recovery, and DR testing artifacts as objective evidence supporting the enterprise resiliency objectives owned by the Cybersecurity team. This role provides technical leadership and oversight to GxP and compliance processes governing backup, recovery, and disaster recovery (DR), ensuring applicable regulatory requirements are identified and met.
Serve as the primary technical resource for RTO/RPO governance, recovery testing, and restore evidence across enterprise IT and lab technology infrastructure. Partner closely with Cybersecurity, Infrastructure, Cloud, Network, and Scientific Systems teams to ensure critical systems and data can be reliably restored within business-approved recovery objectives.
Data Backup and Recovery Coverage
- Lead and scale the enterprise data backup and recovery program across cloud, on-prem, hybrid, and lab infrastructure, increasing coverage of critical systems, servers, instruments, and data stores over time
- Serve as the primary technical owner and go-to resource for backup and recovery coverage across both enterprise IT infrastructure and lab/scientific systems, ensuring lab systems, instruments, and associated data stores are identified, scoped, and included in backup and DR planning alongside enterprise systems in AWS and Azure
- Design and implement backup architecture across compute, storage, and database platforms, ensuring backup strategies align with data classification, retention, and recovery requirements
- Partners with Infrastructure, Cloud, Network, and Scientific Systems teams to identify systems and user/instrument generated data lacking adequate backup coverage and drive remediation to close gaps
- Establish and maintain SOPs to regularly test backups (restore testing, integrity validation, sampling strategies) to ensure ongoing reliability and resiliency of the backup solution
- Define and track KPIs for backup health (e.g., backup success rate, coverage percentage, restore test pass rate, time to restore) and report progress to leadership and governance forums
- Troubleshoot and resolve complex backup and recovery failures, driving root cause analysis and long-term corrective actions
Disaster Recovery Engineering and Execution
- Define and implement disaster recovery solutions and architecture for business critical (Tier
1) technology infrastructure, systems and services, spanning cloud, on-prem, and hybrid environments - Lead the design and implementation of foundational DR capability where none currently exists, including DR architecture, tooling, runbooks, and recovery procedures for Tier 1 systems and platforms
- Establish critical service tiers in partnership with Cyber, business and technology owners, and recommend and establish corresponding RTO (Recovery Time Objective) and RPO (Recovery Point Objective) targets for each tier.
- Plan and execute DR exercises on a regular, recurring cadence, coordinating cross functional participation across Infrastructure, Network, Cloud, and application teams
- Identify and remediate gaps uncovered during DR testing (architecture, dependencies, documentation, automation) to continuously improve recovery readiness
- Build automation to streamline DR failover/failback processes and reduce manual effort and recovery time during actual DR events
Resiliency Evidence, Reporting, and Cyber Partnership
- Maintain a long-term repository of…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).