×
Register Here to Apply for Jobs or Post Jobs. X

Resiliency Engineer

Job in Georgetown, Scott County, Kentucky, 40324, USA
Listing for: KeyBank
Full Time position
Listed on 2026-09-28
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Systems Engineer, Disaster Recovery IT
Salary/Wage Range or Industry Benchmark: 63000 - 96000 USD Yearly USD 63000.00 96000.00 YEAR
Job Description & How to Apply Below

Location

555 Patroon Creek Boulevard, Albany New York

Position Summary

The Resiliency Engineer designs, builds, and continuously improves the reliability, availability, and recoverability of Key Bank's technology platforms across on-premises, hybrid, and cloud (GCP/Azure) environments. Applying software engineering discipline to operations, this role engineers systems to meet defined recovery objectives, automates recovery and validation, and proves that critical services can withstand and recover from failure.

This is a hands-on engineering role for a builder who can write automation, reason about distributed-system failure modes, and facilitate across application, infrastructure, architecture, cloud, and line-of-business teams. The ideal candidate is equally comfortable in code, in a design review, and in a command center during a live recovery exercise.

Key Responsibilities

Automation & Engineering
  • Design and code automation that reduces operational toil and replaces manual, error-prone runbooks with orchestrated, auditable failover and recovery workflows.
  • Develop and maintain infrastructure-as-code, scripts, and pipelines (e.g., Python, Bash, Power Shell, Terraform, Ansible) to provision, configure, and validate recovery environments.
  • Build self-healing patterns, health checks, and automated validation that confirm recoverability before a disaster is ever declared.
Reliability & Resiliency Design
  • Partner with application and infrastructure teams to assess system architecture for reliability, redundancy, and recoverability against assigned system criticality.
  • Define, measure, and support Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for critical services.
  • Ensure architecture is selected to meet Recovery Time Objective (RTO) and Recovery Point Objective (RPO) targets.
Testing, Chaos & Validation
  • Plan and execute disaster recovery tests and targeted fault-injection / chaos experiments (e.g., zonal failure, load-balancer, regional failover) to proactively expose weaknesses.
  • Improve monitoring, alerting, and observability to detect service degradation early.
Facilitation & Partnership
  • Facilitate resiliency and architecture reviews, tabletop exercises, and cross-team recovery walkthroughs, aligning technical and business stakeholders toward clear outcomes.
  • Provide subject-matter expertise on reliability engineering practices and drive adoption across technology teams.
  • Produce clear, examiner-ready documentation and evidence, ensuring work aligns with Key Bank policies, standards, and regulatory requirements.
Required Qualifications
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field-or equivalent work experience.
  • Demonstrated experience in reliability engineering, Dev Ops, infrastructure, or technology operations.
  • Hands-on coding and automation ability (e.g., Python, Bash, Power Shell) and experience with infrastructure-as-code (e.g., Terraform, Ansible).
  • Working knowledge of BOTH on-premises infrastructure (compute, storage, network, virtualization, databases) AND public cloud platforms (GCP and/or Azure).
  • Understanding of high-availability design: redundancy, replication, failover, and load balancing.
  • Strong facilitation, analytical, problem-solving, and written/verbal communication skills, with the ability to influence across teams.
Preferred Qualifications
  • Experience with site reliability engineering principles and service-level management (SLIs, SLOs, error budgets).
  • Experience with disaster recovery planning, resiliency testing, or chaos / fault-injection engineering (e.g., Google FIT, Gremlin).
  • Familiarity with containers and orchestration (Kubernetes/GKE), CI/CD, and observability tooling (e.g., Dynatrace, Prometheus, Grafana, Splunk).
  • Experience with Service Now (ITOM / Business Continuity Management) or comparable orchestration platforms.
  • Experience in a regulated industry or large, complex enterprise environment; familiarity with FFIEC, NIST SP 800-34/CSF, or ISO 22301.
  • Relevant certifications (e.g., cloud architect/engineer, Kubernetes, Linux, ITIL).
COMPENSATION AND BENEFITS

This position is eligible to earn a base salary in the range of $63,000.00 - $96,000.00 annually. Placement within the pay range may differ based upon various factors, including but not limited to skills, experience and geographic location. Compensation for this role also includes eligibility for incentive compensation which may include production, commission, and/or discretionary…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary