Lead Site Reliability Engineer
Listed on 2026-08-29
-
IT/Tech
SRE/Site Reliability, Systems Engineer
In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Lead Site Reliability Engineer position at Vice President level, which is part of the job family responsible for overseeing the production environment, ensuring the operational reliability of deployed software, and implementing strategies to optimize performance and minimize downtime.
Morgan Stanley is an industry leader in financial services, known for mobilizing capital to help governments, corporations, institutions, and individuals around the world achieve their financial goals.
Interested in joining a team that’s eager to create, innovate and make an impact on the world? Read on.
The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the production systems. This position is focused on user and systems support, answering hotline calls, monitoring systems alerts, and taking corrective action. Technical understanding is important as well as the ability to speak to users and understand their problems. In addition to direct user support tasks, the team performs infrastructure related tasks including process configuration, hardware capacity planning, event management, release work, and support tool development to ensure any repetitive tasks are packaged to remove any element of risk.
This role will be responsible for overall stability of the Wealth Management Investment Management application platforms, participation in key optimization initiatives, and collaboration with multiple technical teams within Morgan Stanley. Partner with WM business units, various levels of management and staff to collect, analyze and make recommendations on optimizing the platform. As a team member with expertise in deep analytical triage, you will provide subject matter expertise in debugging, issue analysis and troubleshooting, working with business and technical colleagues to provide reviews and recommendations to avoid any future application issues.
Whatyou’ll do in the role:
Drive Reliability Engineering Practices
Champion SRE principles, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), Error Budgets, and operational risk management frameworks.
Define and track reliability metrics that measure platform health, customer experience, and operational effectiveness.
Lead initiatives focused on reducing Mean Time to Detect (MTTD), Mean Time to Identify (MTTI), and Mean Time to Restore (MTTR).
Own Production ReliabilityProvide leadership for the proactive detection, triage, and resolution of production issues impacting business-critical applications and services.
Serve as the primary owner for escalated production incidents, driving resolution efforts across application, infrastructure, vendor, and external partner teams until service is restored and client impact is mitigated.
Establish a culture of operational excellence focused on stability, resiliency, and continuous service improvement.
Ensure clear, concise, and timely communication during outages, providing accurate business impact assessments and recovery updates to senior leadership.
Production Governance & Change ManagementServe as a key gatekeeper for the production environment, ensuring adherence to change management policies, release controls, operational readiness standards, and risk management practices.
Assess the operational impact of technology changes and ensure appropriate testing, rollback strategies, monitoring, and support models are in place prior to production deployment.
Partner with development teams throughout the software lifecycle to ensure reliability, observability, and operational supportability are built into new applications and services.
Automation & Operational EfficiencyIdentify opportunities to eliminate manual effort and operational toil through automation, self-healing capabilities, and AI-driven operational workflows.
Lead the development and adoption of automation solutions that improve reliability, reduce risk, and increase operational efficiency across the organization.
Promote a culture of engineering-led operations and continuous process optimization.
Operational Readiness & Knowledge ManagementEstablish and maintain a comprehensive knowledge management framework, ensuring runbooks, troubleshooting guides, standards, and operational procedures are accurate, current, and accessible.
Drive operational readiness programs that improve first-level diagnosis and reduce dependency on development teams for routine issue resolution. Create End-to-End Know your system diagrams.
Ensure support teams maintain high-quality documentation and standardized troubleshooting practices to accelerate incident resolution.
Technical LeadershipAct as a senior technical leader and trusted advisor for reliability, resiliency, observability, and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).