More jobs:
Site Reliability Engineer IV
Job in
Buffalo, Erie County, New York, 14266, USA
Listed on 2026-07-11
Listing for:
M&T Bank
Full Time
position Listed on 2026-07-11
Job specializations:
-
IT/Tech
SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below
Overview
Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence, observability, testing, and automation across the Software Development Lifecycle. Leads complex initiatives, influences enterprise engineering standards, and partners with senior stakeholders to improve system stability, resiliency, performance, and operational maturity.
Serves as a mentor and technical leader for less experienced engineers across Technology.
- Accountable for defining and driving service reliability standards, including SLOs, SLAs, SLIs, and error budgets across platforms.
- Design and implement highly available, fault-tolerant architectures aligned with enterprise scalability and resiliency requirements.
- Lead initiatives to improve system reliability, availability, performance, and operational excellence through automation and engineering best practices.
- Develop and promote observability strategies leveraging logging, monitoring, alerting, distributed tracing, Open Telemetry (OTel), Dynatrace, dashboards, and telemetry analytics.
- Design and maintain end-to-end monitoring solutions that provide actionable insights into application, infrastructure, and customer experience health.
- Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.
- Lead incident management practices, including detection, response, escalation, recovery, and coordination of high‑severity production events.
- Drive problem management and Root Cause Analysis (RCA) activities to prevent systemic issues and ensure corrective actions are implemented.
- Lead automation initiatives for self‑healing systems, operational workflows, deployments, recovery procedures, and reliability controls.
- Partner with development teams to build reliable, observable, and scalable services throughout the Software Development Lifecycle (SDLC).
- Design, develop, and execute automated regression testing strategies to validate application stability, reliability, and performance following deployments and infrastructure changes.
- Review test coverage and reliability validation approaches to ensure comprehensive testing and risk mitigation.
- Create, maintain, and improve Infrastructure as Code (IaC) solutions using Terraform for infrastructure provisioning, configuration management, and environment standardization.
- Support and optimise cloud environments, including Microsoft Azure services, deployment automation, scaling strategies, and application lifecycle management.
- Utilize cloud‑native monitoring and operational tools to improve platform visibility, reliability, and performance.
- Serve as a technical authority for performance engineering, resilience, capacity planning, and workload optimisation.
- Drive production readiness practices, including performance testing, resiliency testing, failover validation, disaster recovery preparedness, and operational readiness reviews.
- Review architectural designs and technical roadmaps, providing recommendations to improve reliability, scalability, resiliency, and operational efficiency.
- Lead cross‑team reliability improvement initiatives and influence enterprise engineering standards.
- Develop and maintain operational runbooks, incident playbooks, knowledge articles, and standard operating procedures.
- Partner with development, infrastructure, cybersecurity, architecture, and support teams to identify risks, drive continuous improvement, and optimise platform performance.
- Participate in and lead post‑incident reviews, ensuring actionable outcomes and measurable improvements.
- Communicate system health, reliability trends, operational risks, and remediation strategies to technical and business stakeholders.
- Present reliability initiatives, operational metrics, and engineering recommendations in architecture reviews, technical forums, and leadership discussions.
- Mentor engineers on reliability engineering, observability, cloud engineering, automation,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×