×
Register Here to Apply for Jobs or Post Jobs. X

Vice President - Lead Site Reliability Engineer

Job in Plano, Collin County, Texas, 75086, USA
Listing for: JPMorgan Chase
Full Time position
Listed on 2026-09-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 150000 - 230000 USD Yearly USD 150000.00 230000.00 YEAR
Job Description & How to Apply Below

Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.

As a Lead Site Reliability Engineer at JPMorgan

Chase within the Infrastructure Platforms(Web Hosting) team, you hold a leadership role on your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business challenges they face. You will lead resiliency design reviews, break complex problems into digestible work for other engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to your peers.

In this role, you will partner with stakeholders across engineering and the business to drive reliability, stability, and continuous improvement across web hosting platforms. Your expertise will directly influence service levels, incident response, and the adoption of modern reliability practices - including the responsible use of enterprise-authorized AI capabilities - ensuring the firm's platforms remain resilient, scalable, and secure.

Job responsibilities
  • Consistently model and champion site reliability culture and practices, documenting and sharing knowledge across your organization through internal forums and communities of practice

  • Lead initiatives to improve the reliability and stability of web hosting platforms using data-driven analytics to improve service levels, proactively identifying and resolving technology-related bottlenecks within your areas of expertise

  • Drive collaboration with your team to identify comprehensive service level indicators and partner with stakeholders to establish reasonable service level objectives and error budgets with customers

  • Serve as the primary point of contact during major incidents for your applications, applying deep technical expertise to identify and resolve issues quickly to minimize business impact

  • Use enterprise-authorized AI capabilities to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data in accordance with sensitivity and security requirements

  • Lead reuse-first adoption of AI-assisted reliability workflows across software development lifecycle and toolchain practices - including continuous integration/continuous delivery quality checks, test and validation automation, and operational readiness - ensuring traceability, auditability, resiliency, and security controls

  • Provide ongoing guidance, tools, and solutions to support the firm's growth while working toward deep expertise on the applications and platforms within your scope, including their interdependencies and limitations

  • Offer mentorship and technical advice to other engineers, helping to grow site reliability knowledge and capability across the team

Required qualifications, capabilities, and skills
  • Formal training or certification on site reliability engineering concepts and 5+ years applied experience

  • Demonstrated proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and other site reliability best practices, with the ability to implement these within an application or platform

  • Advanced knowledge of site reliability culture and principles with demonstrated ability to apply them within an application or platform environment

  • Proficient knowledge and experience in observability, including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Cloud Watch, or Splunk

  • Demonstrated experience using enterprise-authorized AI capabilities to improve site reliability engineering workflows - such as incident investigation support and knowledge capture - with strong validation habits and awareness of data sensitivity

  • Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations

  • Fluency in at least one programming language such as Python, Java/Spring Boot, or Ansible

  • Proficient with continuous…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary