Lead Site Reliability Engineer – Operations Excellence
Listed on 2026-08-17
-
IT/Tech
SRE/Site Reliability
Job Description
Join a team where your engineering expertise directly protects the reliability and resilience of systems that matter JPMorgan
Chase, we invest in engineers who think beyond the code — who own outcomes, drive operational excellence, and raise the bar for how technology performs under pressure. As a Lead Software Engineer focused on Site Reliability and Operations Excellence, you will have the opportunity to shape how we respond to, learn from, and prevent incidents across a complex, high-stakes technology environment.
This is a role where your impact is visible, your voice carries weight, and your work directly influences the stability of services relied upon by millions.
As a Lead Software Engineer - Site Reliability Engineer, Operations Excellence at JPMorgan
Chase within the AI/ML & Data Platforms area, you will serve as a technical leader at the intersection of software engineering and reliability engineering — owning the incident management lifecycle, defining reliability standards, and partnering across Engineering, Product, Infrastructure, and Security to deliver durable operational improvements. You will bring structure to complexity, clarity to high-urgency situations, and a continuous improvement mindset to everything from alert quality to executive reporting.
Your work will directly strengthen the firm's ability to detect, respond to, and prevent production issues at scale.
- Own and continuously improve the incident management lifecycle, including triage, escalation, stakeholder communications, and recovery, ensuring consistent execution and measurable improvement over time.
- Serve as Incident Commander for major incidents, coaching responders to follow defined processes and site reliability engineering best practices while maintaining clear and timely stakeholder communication.
- Drive operational readiness through drills, game days, and failure-mode exercises that improve response effectiveness, surface gaps, and build team resilience before incidents occur.
- Define, implement, and enforce reliability standards across services, including service level indicators and objectives, error budgets, monitoring coverage, and alert quality.
- Strengthen change and release management practices by establishing readiness checks, progressive delivery standards, rollback procedures, runbook quality, and production hygiene expectations.
- Partner with Engineering, Product, Infrastructure, and Security teams to prioritize reliability work and deliver solutions that address user and operational pain points with lasting impact.
- Lead problem management and root cause analysis by debugging complex production issues, identifying systemic root causes, and driving durable remediation and preventative actions.
- Own reliability reporting and operational governance by tracking key performance indicators — including availability versus service level objectives, mean time to detect and recover, incident trends, alert noise, and change failure rate — and producing executive-ready summaries.
- Facilitate operational forums including operations reviews, incident review boards, and reliability councils to drive accountability, share learnings, and align stakeholders on priorities.
- Drives team adoption of enterprise-authorized AI-assisted engineering practices within the work environment to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test strategy acceleration, incident/root-cause analysis support), while establishing consistent validation standards (secure coding, peer review, automated testing) and promoting reuse of effective patterns across the team.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
- Formal training or certification on software engineering concepts and advanced applied experience.
- Demonstrated experience improving operational processes, including incident management, root cause analysis, problem management, and…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: