Site Reliability Engineering; SRE ) lead
Listed on 2026-09-12
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Support
At U.S. Bank, we’re on a journey to do our best. Helping the customers and businesses we serve to make better and smarter financial decisions and enabling the communities we support to grow and succeed. We believe it takes all of us to bring our shared ambition to life, and each person is unique in their potential. A career with U.S. Bank gives you a wide, ever-growing range of opportunities to discover what makes you thrive at every stage of your career.
Try new things, learn new skills and discover what you excel at—all from Day One.
Lead the troubleshooting and resolution of complex production incidents
, including application failures, API issues, cloud platform outages, performance degradation, and operational disruptions.Conduct comprehensive root cause analysis (RCA), impact assessments, mitigation planning, and implementation of permanent corrective actions.
Design and enhance monitoring, observability, alerting, dashboards, health checks, and operational runbooks to improve platform reliability and availability.
Drive automation initiatives using scripting, Infrastructure as Code (IaC), CI/CD pipelines, and self-healing capabilities to reduce manual operational effort.
Partner with software engineering, infrastructure, and product teams to identify, prioritize, and remediate recurring reliability issues.
Serve as the Incident Commander during major incidents, coordinating cross-functional response teams and driving restoration activities.
Provide leadership, coaching, mentoring, and workload management for SRE, Dev Ops, and production support engineers.
Utilize operational metrics including MTTR, MTTD, SLA compliance, backlog health, incident volume, and problem closure rates to drive continuous improvement and operational excellence.
- Bachelor's degree, or equivalent work experience
- Six to eight years of relevant work experience in business and risk analysis, IT Service Management, production support, product/project management, or application development
- Strong expertise in Site Reliability Engineering (SRE), Dev Ops, Production Support, Platform Engineering, and Distributed Systems Operations.
- Experience leading technical teams, incident response efforts, workload prioritization, and reliability improvement programs.
- Advanced knowledge of Incident Management, Problem Management, Change Management, and Root Cause Analysis (RCA) methodologies.
- Hands-on experience with AWS, Azure, Kubernetes, Docker, and cloud-native infrastructure platforms.
- Proficiency with Python, Power Shell, Shell Scripting, and automation frameworks for operational efficiency and reliability engineering.
- Experience building and supporting CI/CD pipelines using tools such as Git Hub Actions, Azure Dev Ops, Jenkins, or Git Lab.
- Strong expertise in Monitoring and Observability Solutions including Datadog, Splunk, Dynatrace, Grafana, Prometheus, Cloud Watch, Azure Monitor, and Open Telemetry.
- Experience with Service Now, Jira, Terraform, Ansible, REST APIs, SQL/Relational Databases
, along with excellent stakeholder communication and leadership skills.
- AWS Certified Solutions Architect, Dev Ops Engineer, or equivalent AWS certification
- Microsoft Azure Administrator, Architect, or Dev Ops Engineer certification
- Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD)
This role requires working from a U.S. Bank location three (3) or more days per week.
If there’s anything we can do to accommodate a disability during any portion of the application or hiring process, please refer to disability accommodations for applicants.
Benefits:- Healthcare (medical, dental, vision)
- Basic term and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).