Senior Site Reliability Engineer
Job in
Wayne, Delaware County, Pennsylvania, 19080, USA
Listed on 2026-08-05
Listing for:
Vanguard
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
In addition, you will automate incident response capabilities and pioneer AI‑enhanced diagnostics and analysis to improve detection, response, and recovery. You will work alongside a collaborative, technically focused team where your innovations in resiliency engineering directly shape Vanguard's next generation of reliable, client‑centric experiences.
Core Responsibilities:
* Evaluate applications, platforms, and vendors to assess resiliency, reliability, and operational risk.
* Design and implement processes that enforce enterprise resiliency and reliability standards.
* Lead blameless post‑incident reviews for high‑severity incidents or incidents spanning multiple complex product families.
* Partner with product and platform teams to proactively identify and remediate reliability risks before they impact clients.
* Develop, communicate, and evangelize new standards, tools, and frameworks across subdivisions, ensuring consistent adoption.
* Troubleshoot complex production issues and implement durable solutions that prevent recurrence.
* Participate in a periodic on‑call rotation to support production stability.
* Evaluate and onboard resiliency and reliability tooling.
* Actively participate in reliability engineering and resilience communities of practice, contributing to shared learning and enterprise consistency.
* Contribute to strategic initiatives that advance Vanguard's operational maturity and resiliency posture.
Qualifications | Technical
Skills:
* Observability Platforms:
Experience with modern observability and monitoring tools, such as Splunk, Honeycomb, Cloud Watch, Dynatrace, or App Dynamics.
* Reliability Metrics:
Strong understanding of SLIs, SLOs, and SLAs, including dashboarding and reporting practices.
* Monitoring & Alerting:
Experience with alert design, anomaly detection, predictive alerting, and synthetic monitoring using structured methodologies.
* Automation & Resilience Engineering:
Experience with automation and resilience practices such as Python-based automation, RPA platforms (e.g., Blue Prism, UiPath), chaos engineering, and failure analysis techniques (e.g., FMEA).
Special Factors
Sponsorship
Vanguard is not offering visa sponsorship for this position.
About Vanguard
At Vanguard, we don't just have a mission-we're on a mission.
To work for the long-term financial wellbeing of our clients. To lead through product and services that transform our clients' lives. To learn and develop our skills as individuals and as a team. From Malvern to Melbourne, our mission drives us forward and inspires us to be our best.
How We Work
Vanguard has implemented a hybrid working model for the majority of our crew members, designed to capture the benefits of enhanced flexibility while enabling in-person learning, collaboration, and connection. We believe our mission-driven and highly collaborative culture is a critical enabler to support long-term client outcomes and enrich the employee experience.
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×