Lead Site Reliability Engineer
Job in
Buffalo, Erie County, New York, 14266, USA
Listed on 2026-08-02
Listing for:
M&T Bank Corporation
Full Time
position Listed on 2026-08-02
Job specializations:
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Primary Responsibilities
- Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices.
- Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence.
- Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.
- Develop comprehensive observability strategies leveraging Dynatrace, Open Telemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions.
- Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.
- Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.
- Lead incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer impact.
- Perform and facilitate Root Cause Analysis (RCA) activities, ensuring corrective and preventive actions are identified, prioritized, and implemented.
- Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
- Partner with development teams to build reliable and observable services throughout the Software Development Lifecycle (SDLC).
- Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.
- Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.
- Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for cloud infrastructure provisioning, configuration management, and environment standardization.
- Support and optimize Microsoft Azure environments, including Azure App Services, resource management, scaling strategies, deployment automation, and application lifecycle management.
- Utilize Azure-native tools such as Azure Monitor, Application Insights, Log Analytics, and related services to improve platform visibility and reliability.
- Drive implementation of performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness within assigned domains.
- Establish operational readiness standards and ensure applications meet reliability, scalability, observability, and supportability requirements before production deployment.
- Review architectural designs and provide recommendations to improve platform resiliency, operational efficiency, and cloud optimization.
- Lead capacity planning, performance tuning, and workload optimization efforts across production environments.
- Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.
- Serve as a key partner with engineering, infrastructure, cybersecurity, architecture, and support teams to identify and implement continuous process improvements spanning organizational boundaries.
- Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
- Present reliability initiatives, operational metrics, and engineering recommendations at architecture reviews, technical forums, and leadership meetings.
- Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices.
- Understand and adhere to the Company's risk and regulatory standards, policies, and controls in accordance with the Company's Risk Appetite.
- Identify reliability, operational, and technology risks requiring escalation to management.
- Promote an environment that supports a culture of belonging and reflects the M&T Bank brand.
- Maintain M&T internal control standards, including timely implementation of internal and external audit findings and regulatory requirements as applicable.
- Complete other related duties as assigned.
Supervisory/Managerial Responsibilities
No supervisory responsibilities. May provide technical leadership, mentorship, and guidance to engineers and project teams.
Education and Experience RequiredAssociate's degree and a minimum of 7 years' systems analysis and/or application development work experience or Bachelor's degree and a minimum of 5 years' systems analysis and/or application development work experience. In lieu of a degree, a combined minimum of 9 year's education and/or relevant work experience, including a minimum of 5 years' system analysis and/or application development work experience.
Core Requirements- Strong experience in observability and monitoring, including hands-on expertise with:
- Dynatrace
- Open Telemetry (OTel)
- Distributed tracing
- Metrics collection and analysis
- Centralized logging and log aggregation
- Alerting and…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×