Senior Lead Technology Resiliency Engineer
Job in
Tucson, Pima County, Arizona, 85718, USA
Listed on 2026-08-16
Listing for:
Jobtailor
Full Time
position Listed on 2026-08-16
Job specializations:
-
IT/Tech
Disaster Recovery IT, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
- Define and evolve the enterprise technology resiliency strategy, aligned to business critical services, regulatory expectations, and long‑term technology roadmaps
- Establish resiliency design principles and patterns for cloud, hybrid, and on‑prem platforms (e.g., multi‑region, multi‑AZ, active/active, degradation strategies)
- Influence architecture decisions to ensure resiliency, recoverability, and operability are built in from inception—not retrofitted
- Partner with enterprise and domain architects to embed resiliency requirements into standards, reference architectures, and engineering practices
- Support identification and mapping of Important Business Services and underpinning technology dependencies
- Translate business impact tolerances into actionable technology recovery objectives (RTO, RPO, MTO)
- Assess end‑to‑end service resilience, identifying single points of failure across applications, data, infrastructure, vendors, and people
- Drive remediation strategies for material resiliency gaps, balancing risk reduction with real‑world delivery constraints
- Act as a trusted advisor to engineering teams on fault tolerance, high availability, disaster recovery, and graceful degradation
- Review major platform and system designs from a resiliency and operability perspective
- Promote modern resiliency practices such as:
Chaos engineering and failure injection, Automation‑first recovery, Observability and service‑level indicators, Immutable infrastructure and infrastructure‑as‑code - Shape the firm’s approach to resilience testing, including scenario‑based testing, disaster recovery exercises, and severe‑but‑plausible events
- Design and participate in enterprise simulation exercises involving technology and business stakeholders
- Drive a culture of continuous learning through post‑incident analysis focused on systemic improvement rather than blame
- Ensure lessons learned feed directly into architecture, standards, and engineering practices
- Partner with Risk, Compliance, and Audit to ensure resiliency practices meet internal policy and external regulatory expectations
- Contribute to regulatory responses, exams, and remediation programs related to operational and technology resilience
- Help define meaningful, decision‑useful resiliency metrics and management reporting
- 7+ years of Systems Engineering, Technology Architecture experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
- 7+ years of experience in large‑scale technology engineering, architecture, SRE, platform engineering, or infrastructure roles
- 7 plus years Proven experience designing and operating highly available, distributed systems in complex enterprise environments
- 7+ years' experience with:
Cloud and hybrid architectures (AWS, Azure, GCP or equivalent) - Disaster recovery and availability patterns
- Data replication and consistency models
- Observability, monitoring, and incident management
- 7+ years' experience influencing architecture and design decisions at scale
Demonstrates expertise in defining and evolving enterprise technology resiliency strategies, ensuring alignment with business critical services and regulatory expectations. Proficient in designing and implementing cloud and hybrid architectures while promoting modern resiliency practices and continuous improvement.
Highest‑signal resume keywords- Enterprise Technology Resiliency Strategy
- Cloud And Hybrid Architectures
- Disaster Recovery And Availability Patterns
- Observability And Incident Management
- Influencing Architecture Decisions
- Systems Engineering
- Technology Architecture
- Distributed Systems Design
- Data Replication Models
- Resiliency Metrics
- Recovery Objectives (RTO, RPO, MTO)
- Chaos Engineering
- Automation-First Recovery
- Infrastructure-As-Code
- Resilience Testing
- Trusted Advisor
- Continuous Learning Culture
- Collaboration With Stakeholders
- Regulatory Compliance
- Operational Resilience
- Risk Management
- Business Impact Tolerances
- Engineering Practices
- AWS
- Azure
- GCP
- Monitoring Tools
- Incident Management Systems
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×