Sr. Cloud Resilience Engineer - Security
Listed on 2026-07-22
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Architectural Leadership & Consulting
Act as the primary resilience advisor to multiple distributed product and enterprise teams, guiding them on best practices for building high availability (HA) and redundancy into their SaaS applications.
Resilient Cloud DesignDesign and recommend fast-failover solutions and highly available infrastructure primarily on Google Cloud Platform (GCP), while also providing oversight for workloads in Azure and AWS.
Infrastructure ValidationLeverage your strong background in Infrastructure as Code (IaC) to review, validate, and guide the implementation efforts of engineering teams.
Container & Compute ResilienceDesign redundancy strategies for workloads running on Google Kubernetes Engine (GKE) and virtual machines, ensuring self‑healing deployments.
Cross‑Functional CollaborationPartner closely with Dev Ops, SRE, and Product Engineering teams to champion resilience engineering principles, chaos testing, and failover validations across tier‑0 mission‑critical systems.
Cloud Platform ExpertiseDeep, practical technical knowledge of Google Cloud Platform (GCP) core services, specifically GKE, Compute Engine, and CloudSQL. Familiarity with AWS and Azure is highly desirable.
Technical Practitioner BackgroundProven past experience as a hands‑on engineer who has deployed complex infrastructure. You should understand the implementation details well enough to effectively guide engineering teams.
High Availability ArchitectureDemonstrated success in architecting active‑active or active‑passive fast‑failover mechanisms for high‑volume, data‑intensive SaaS applications.
Database ResilienceStrong understanding of database clustering, replication, and migration strategies (especially migrating legacy RDBMS like MS SQL Server to cloud‑native solutions like CloudSQL).
Advisory SkillsExcellent communication and consulting skills, with the ability to influence technical teams, explain complex architectural concepts, and foster a culture of resilience without having direct reporting authority over the engineering teams.
Resilient Network ServicesPractical design experience managing high‑availability network topologies, including load balancing, DNS & name resolution, firewalls/gateways, identity/authentication systems, and centralized logging/SIEM.
User Session ManagementDeep understanding of user session replication, session state persistence, and failover routing strategies in high‑traffic, multi‑region application architectures.
Broad HA Domain ExposureFamiliarity assessing or designing resilience across a comprehensive range of critical SaaS failure domains, such as API gateways, caching layers, messaging/queuing systems, and CI/CD pipelines.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).