Journey-Centric Lead Site Reliability Engineer; SRE
Listed on 2026-09-13
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Journey-Centric Lead Site Reliability Engineer (SRE)
Phoenix, AZ (Hybrid)
JPC - 20640
We are seeking a highly experienced Journey-Centric Lead Site Reliability Engineer (SRE) with Banking and Financial Services (BFS) domain experience to drive end-to-end reliability, observability, automation, and operational excellence across critical customer and business journeys. This role will lead to the design and implementation of modern SRE practices, unified observability platforms, self-healing capabilities, AI-driven operations, and workflow automation to ensure highly resilient, scalable, and intelligent digital services.
The ideal candidate combines deep SRE expertise with strong platform engineering, cloud operations, networking, observability, and AI Ops experience, with a particular focus on AWS Agent Core-powered operational intelligence and autonomous operations.
Qualifications:
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, Cloud Operations, or Dev Ops with Banking and Financial Services customers.
- 3+ years leading enterprise-scale reliability transformation initiatives.
- Experience managing mission-critical digital platforms and customer journeys.
- Experience with production support and handling critical SLOs.
- Strong knowledge of:
- SRE principles and practices
- Reliability engineering frameworks
- Distributed systems
- Cloud-native platforms
- Hands-on experience with one or more:
- Dynatrace
- New Relic
- Splunk Observability
- Elastic Stack
- Open Telemetry
- Automation & Engineering experience with:
- Terraform
- Cloud Formation
- Ansible
- Git Hub Actions
- Jenkins
- Python
- Go
- Bash/Shell scripting
- AI Ops & GenAI
- Experience implementing AI Ops solutions. - Knowledge of autonomous operations and intelligent remediation systems.
- Hands-on experience with AWS Agent Core and agent-based operational platforms.
- Familiarity with AI-powered observability and incident management tools.
- Networking Knowledge
- Strong understanding of Network Layer: - TCP/IP
- DNS
- HTTP/HTTPS
- API Gateways
- Load Balancers
- CDN
- Service Mesh
- Network Performance Engineering
Preferred Qualifications (Desired)
- Experience in customer journey monitoring and digital experience.
Responsibilities:
SRE & Reliability Engineering
- Define and implement enterprise-scale SRE best practices across critical applications and digital journeys.
- Establish reliability frameworks, operational standards, and governance models.
- Drive proactive reliability engineering initiatives to improve system availability, resilience, and performance.
- Lead incident management, postmortem analysis, root cause investigations, and reliability reviews.
Unified Observability & Monitoring
- Design and implement a unified observability strategy encompassing metrics, logs, traces, events, and user experience telemetry.
- Build comprehensive observability dashboards for business and technology stakeholders.
- Implement distributed tracing and end-to-end monitoring across complex microservices ecosystems.
- Define observability standards and instrumentation frameworks across engineering teams.
- Enable unified logs, metrics, and trace correlation capabilities for rapid issue detection and troubleshooting.
- Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through observability-driven insights.
- Establish service dependency mapping and journey-centric operational visibility.
Self-Healing & Autonomous Operations
- Design and implement self-healing capabilities using event-driven automation and AI-assisted remediation.
- Develop automated recovery processes for common failure scenarios.
- Create autonomous operational workflows that minimize manual intervention.
- Integrate predictive alerting and automated response mechanisms.
Automation & Workflow Engineering
- Build scalable operational automation frameworks.
- Develop infrastructure, application, observability, and operational workflows using Infrastructure as Code (IaC), Monitoring as Code (MaC), and Observability as Code (OaC).
- Automate deployments, monitoring, remediation, and operational runbooks.
- Reduce operational toil through intelligent engineering solutions.
SLO, SLA & Error Budget Management
- Define and govern measurable Service Level Objectives (SLOs), Service Level Agreements (SLAs), and Error Budgets.
- Partner with engineering and business teams to align reliability targets with customer expectations.
- Establish service maturity metrics and reliability scorecards.
- Drive data-driven operational decision-making through reliability KPIs.
Network & Platform Reliability
Apply deep understanding of:
- Understanding of network…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).