×
Register Here to Apply for Jobs or Post Jobs. X

Engineer, Site Reliability

Job in Charlotte, Mecklenburg County, North Carolina, 28201, USA
Listing for: Vanguard
Full Time position
Listed on 2026-09-01
Job specializations:
  • Software Development
    Cloud Engineer - Software, DevOps, Backend Developer, Software Architect
Job Description & How to Apply Below

hackajob is collaborating with Vanguard to connect them with exceptional professionals for this role.

Join the Personal Wealth Technology (PWTech) Reliability Engineering team and lead cutting-edge Reliability Engineering initiatives that impact hundreds of applications and millions of investors. You'll architect and build enterprise-scale resiliency solutions, driving our ambitious roadmap. This is an opportunity to combine deep technical expertise with strategic influence
-automating incident responses, implementing distributed tracing at scale, and pioneering AI-enhanced diagnostics and analysis. Work alongside a collaborative, technically-focused team where your innovation in resilience engineering will shape Vanguard's next generation of client experiences.
At Vanguard, we pride ourselves on delivering an exceptional client experience to all investors; at the core of this experience are systems that reside in a technically complex and constantly evolving resiliency landscape. Passionate, technically skilled engineers are at the center of our resiliency operations, and we are looking to grow our team.
We are seeking an experienced engineer with broad, end-to-end software development experience, including operating applications in a microservices environment in production s role goes beyond feature implementation - it requires someone who can design, build, and support resilient systems from the ground up.
As a Staff Reliability Engineer at Vanguard, you will play a critical role in solving impactful operational problems. You are curious and take a proactive approach to identifying problems and making improvements. You balance innovative thinking with pragmatism and understand the long-term impacts of technical decisions. You communicate complex ideas clearly and collaborate effectively to deliver scalable solutions.

Core Responsibilities

  • Lead the technical strategy, architecture, and evolution of PWTech reliability engineering platforms and capabilities, ensuring they scale across hundreds of applications and critical client-facing systems.
  • Design and build production-grade software and platforms that improve reliability outcomes, including automated incident detection, diagnostics, remediation, resiliency engineering, and operational intelligence.
  • Drive enterprise observability and diagnostics capabilities, enabling consistent telemetry, distributed tracing, metrics, and operational insights across cloud-native technologies and application stacks.
  • Define and codify resilient application and platform patterns, such as graceful degradation, circuit breakers, load shedding, fault isolation, failover, and automated recovery, driving adoption through reusable software, frameworks, and engineering standards.
  • Influence engineering teams and technology leaders across the organization, establishing technical standards and ensuring reliability is designed into systems from inception.
  • Lead the resolution of complex reliability and production challenges, identifying systemic risks, driving root cause analysis, and engineering durable solutions that improve long-term resilience.
  • Participates in special projects and performs other duties as assigned.

Qualifications

  • Minimum of eight years related experience, with at least two years of development experience.
  • Undergraduate degree or equivalent combination of training and experience. Graduate degree preferred.

Preferred Skills

  • Experience designing, building, and operating production-facing platforms or engineering capabilities that achieve broad adoption and deliver measurable reliability, operational, or business outcomes.
  • Deep expertise in distributed systems architecture, including scalability, availability, resiliency, fault tolerance, performance optimization, and production operations at scale.
  • Strong technical leadership and influence skills, with a demonstrated ability to drive architecture decisions, establish technical standards, and align multiple teams on engineering direction.
  • Deep expertise in Java or JavaScript, with hands-on experience developing and operating software in modern cloud-native and microservices environments.
  • Demonstrated ability to diagnose…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary