Sr. Availability & Reliability Engineering Manager
Listed on 2026-09-12
-
Software Development
Need Help?
Regular or Temporary
Regular
Language FluencyEnglish (Required)
Work Shift1st shift (United States of America)
Job DescriptionLeads a small to midsized software engineering team, implementing and helping shape key elements of tactical and operational plans established by executive and senior leadership. Drives timely delivery of complex, scalable software solutions for CTO within defined scope and quality parameters, making informed day to day decisions to keep work on track.
The Sr. Availability & Reliability Engineering Manager will lead a software engineering team responsible for building and advancing Availability & Reliability Engineering capabilities that improve system resiliency, reduce incident resolution times, and enhance operational excellence. Drives the development of tooling, observability solutions, automation capabilities, and operational tooling for CTO while partnering with engineering, infrastructure, cloud, and architecture teams to identify systemic issues and improve technology health.
Provides leadership, coaching, and talent development for engineering teammates, fostering a culture of accountability, continuous improvement, operational ownership, and engineering excellence. Translates strategic objectives into executable roadmaps that strengthen platform reliability, enable AI-driven operations, and accelerate delivery of resilient technology solutions.
Truist will not sponsor an applicant for work visa status or employment authorization, nor will we offer any immigration-related support for this position (including, but not limited to H-1B, F-1 OPT, F-1 STEM OPT, F-1 CPT, J-1, TN-1 or TN-2, E-3, O-1, or future sponsorship for U.S. lawful permanent residence status.)
Essential Duties And Responsibilities- Implements and helps refine tactical and operational plans for a software engineering team, translating strategies from senior leaders into clear objectives and daytoday priorities for a small to midsized group.
- Manages lower level managers and experienced individual contributors responsible for designing, developing, testing, and deploying software solutions, ensuring work is delivered on time and to agreed scope and quality.
- Applies and improves technology standards, governance practices, and engineering methods within the team, making moderate to significant enhancements to processes and ways of working.
- Provides data and feedback to senior leadership on team performance and improvement opportunities, influencing how operational plans are adjusted and executed.
- Develops and coaches managers and senior individual contributors, setting expectations, providing regular feedback, and reinforcing a collaborative, results oriented engineering culture.
- Leads process, tooling, and system improvements within the team, using metrics and retrospectives to streamline workflows, reduce waste, and increase reliability of delivery.
- Ensures effective risk management and adherence to regulatory and policy requirements for software development and deployment in the team’s scope, coordinating with partner teams to resolve issues within defined guidelines.
Required Qualifications
- Bachelor’s degree in Computer Science, Software Engineering, or related field.
- Minimum of 7 years of professional experience in software development, with significant leadership experience.
- Advanced management and leadership experience leading teams.
- Developing expertise and knowledge of software architecture, development lifecycle, and technology strategy.
- Advanced degree in Computer Science, Software Engineering, Information Systems, or related discipline.
- Demonstrated success leading engineering teams responsible for platform reliability, observability, site reliability engineering (SRE), cloud engineering, technology operations, or operational resilience initiatives.
- Experience developing and implementing observability strategies utilizing telemetry, logging, monitoring, distributed tracing, and operational analytics to improve service health and reduce Mean Time to Resolution (MTTR).
- Proven ability to lead root cause analysis, incident management, fault isolation, resiliency validation, and restoration strategy activities within complex distributed systems environments.
- Experience building or leveraging AI-enabled operations, automation platforms, intelligent incident response capabilities, anomaly detection, change correlation, or agentic AI solutions to improve operational…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).