×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer Lead

Job in Plano, Collin County, Texas, 75023, USA
Listing for: Bank of America
Full Time position
Listed on 2026-07-29
Job specializations:
  • IT/Tech
    Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Support
Job Description & How to Apply Below

Site Reliability Engineering (SRE) Leader

At Bank of America, we are guided by a common purpose to help make financial lives better through the power of every connection. We do this by driving Responsible Growth and delivering for our clients, teammates, communities, and shareholders every day. Being a Great Place to Work is core to how we drive Responsible Growth. This includes our commitment to being an inclusive workplace, attracting and developing exceptional talent, supporting our teammates' physical, emotional, and financial wellness, recognizing and rewarding performance, and how we make an impact in the communities we serve.

Bank of America is committed to an in-office culture with specific requirements for office-based attendance and which allows for an appropriate level of flexibility for our teammates and businesses based on role-specific considerations. At Bank of America, you can build a successful career with opportunities to learn, grow, and make an impact. Join us!

This job is responsible for building and leading a team to deliver technology products and services that meet business outcomes. Key responsibilities include developing a technology strategy, ensuring technology solutions comply with applicable standards, promoting design, engineering, and organizational practices, and advocating and advancing modern, Agile solution delivery practices. Job expectations may include coaching, mentoring, providing feedback and hands on career development, identifying emerging talent, fostering leadership skills, and managing stakeholders.

Seeking a seasoned Site Reliability Engineering (SRE) Leader to drive the reliability, scalability, and performance of critical Infrastructure Automation platforms. This role will lead the design and implementation of SRE practices across a federated technology ecosystem, ensuring operational excellence through automation, observability, and resilient architecture.

The ideal candidate will bring deep expertise in distributed systems, cloud-native infrastructure, SaaS application support and Dev Ops/SRE principles, along with strong leadership and collaboration skills to influence cross-functional engineering and Production management teams and drive continuous improvement in service reliability.

Responsibilities
  • SRE Strategy & Governance:
    • Define and implement SRE frameworks, including SLIs/SLOs/SLAs, error budgets, and incident response protocols.
    • Establish governance models for reliability engineering across distributed teams.
    • Champion a culture of observability, proactive monitoring, and continuous feedback loops.
  • Reactive & Proactive Problem Management:
    • Lead root cause analysis (RCA) and post-incident reviews to identify systemic issues and prevent recurrence.
    • Implement proactive problem detection using telemetry, anomaly detection, and trend analysis.
    • Collaborate with engineering and operations teams to eliminate toil and reduce incident frequency and impact.
  • Capacity & Performance Management:
    • Develop and maintain capacity models to ensure systems scale efficiently with business demand.
    • Monitor performance trends and lead optimization efforts across infrastructure and applications.
    • Partner with finance and engineering teams to align capacity planning with cost and growth objectives.
  • Platform Reliability & Automation:
    • Drive automation of operational tasks including deployments, scaling, and recovery.
    • Integrate reliability tooling with CI/CD pipelines, ITSM platforms (e.g., Service Now), and observability systems.
  • Incident Management & Operational Excellence:
    • Oversee major incident response, escalation, and communication processes.
    • Develop and maintain runbooks, playbooks, and escalation protocols.
    • Drive continuous improvement through blameless retrospectives and operational reviews.
  • Technical Leadership:
    • Serve as a senior technical advisor and thought leader in SRE and platform engineering.
    • Mentor and guide SRE teams and partner with engineering leaders across the enterprise.
    • Provide input on staffing, tooling strategy, and budget planning for reliability initiatives.
  • Managerial Responsibilities:
    This position may also have responsibilities for managing associates. At Bank…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary