SRE/Backup Engineer
Listed on 2026-09-07
-
IT/Tech
Disaster Recovery IT, Systems Engineer, Cloud Computing: Infrastructure & Operations, Cybersecurity
Role: SRE/Backup Engineer
Location:
Chicago, IL Duration: 9+ Months Project Overview / Contractor's Role: SRE/Back Up Engineer
We are seeking a highly technical Senior Site Reliability Engineer (SRE) with deep expertise in enterprise backup engineering, cyber recovery, and platform resiliency. This role will be responsible for engineering highly available, secure, and automated recovery capabilities that protect the organization against operational failures, ransomware, and other cyber threats.
The ideal candidate combines traditional SRE principles-automation, observability, reliability engineering, and resilience-with extensive experience designing and operating enterprise backup platforms, immutable storage, air-gapped cyber vaults, isolated recovery environments (IREs), and recovery orchestration. This individual will partner closely with Infrastructure, Cyber Security, Cloud Engineering, Application Development, and Disaster Recovery teams to ensure critical services remain recoverable, resilient, and continuously validated.
Experience Level: 3 - SeniorRequired Qualifications
- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent experience.
- 7+ years in Backup Engineering, Infrastructure Engineering, or Site Reliability Engineering.
- 5+ years designing enterprise backup solutions.
- 3+ years supporting cyber recovery architectures.
- Experience implementing SRE principles within enterprise infrastructure environments.
- Strong understanding of distributed systems and high availability architectures.
- Cohesity
- Dell Power Protect Data Manager
- Dell Data Domain
- Dell Cyber Recovery
- Rubrik
- Commvault
- Veritas Net Backup
- Veeam
- Air-gapped vaults
- Immutable backups
- Clean Rooms
- Isolated Recovery Environments (IRE)
- Recovery orchestration
- Cyber resilience testing
- Ransomware recovery
- Recovery validation
- Microsoft Azure
- AWS
- Google Cloud Platform
- Cloud-native backup
- Cross-region recovery
- Hybrid cloud resiliency
- VMware
- Hyper-V
- Kubernetes
- Open Shift
- Linux
- Windows Server
- Active Directory
- Enterprise storage platforms
- Ansible
- Terraform
- Python
- Power Shell
- Bash
- Git Hub
- Git Hub Actions
- CI/CD pipelines
- Dynatrace
- Grafana
- Prometheus
- Splunk
- ELK Stack
- Service Now
- Zero Trust architecture
- NIST Cybersecurity Framework
- CIS Controls
- Encryption and key management
- Identity and Access Management (IAM)
- Multi-factor authentication (MFA)
- Secure recovery processes
- Experience in financial services or another highly regulated industry.
- Experience supporting GSIB cyber resiliency programs.
- Knowledge of regulatory expectations from agencies such as the Federal Reserve, OCC, or FFIEC.
- Experience with chaos engineering and resilience testing.
- Familiarity with SRE tooling and reliability metrics.
- Experience implementing AI-assisted operations (AIOps) and predictive analytics.
- Strong systems thinking and engineering mindset.
- Excellent troubleshooting and root cause analysis skills.
- Ability to lead cross-functional technical recovery efforts.
- Strong communication and executive presentation skills.
- Proven ability to influence engineering standards and drive operational excellence.
- Commitment to continuous improvement through automation and reliability engineering.
- Engineer and maintain highly available, resilient enterprise platforms using SRE principles.
- Define and measure Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for backup and recovery services.
- Develop automation to reduce operational toil and improve reliability.
- Perform root cause analysis (RCA) and implement permanent corrective actions.
- Continuously improve platform reliability, scalability, performance, and recoverability.
- Establish proactive monitoring, alerting, and observability for backup and cyber recovery platforms.
- Participate in incident response and major incident recovery activities.
- Design, implement, and administer enterprise backup and recovery solutions across on‑premises, cloud, and SaaS platforms.
- Engineer immutable backup…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).