×
Register Here to Apply for Jobs or Post Jobs. X

Lead Site Reliability Engineer

Job in Alpharetta, Fulton County, Georgia, 30022, USA
Listing for: Morgan Stanley
Full Time position
Listed on 2026-08-28
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below

Lead Site Reliability Engineer

In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This is a Lead Site Reliability Engineer position at Vice President level, which is part of the job family responsible for overseeing the production environment, ensuring the operational reliability of deployed software, and implementing strategies to optimize performance and minimize downtime.

Morgan Stanley is an industry leader in financial services, known for mobilizing capital to help governments, corporations, institutions, and individuals around the world achieve their financial goals.

Interested in joining a team that's eager to create, innovate and make an impact on the world? Read on.

The Reliability Operations (RO) within WMT is responsible for providing swift, courteous, and knowledgeable customer service to end users of the production systems. This position is focused on user and systems support, answering hotline calls, monitoring systems alerts, and taking corrective action. Technical understanding is important as well as the ability to speak to users and understand their problems. In addition to direct user support tasks, the team performs infrastructure related tasks including process configuration, hardware capacity planning, event management, release work, and support tool development to ensure any repetitive tasks are packaged to remove any element of risk.

This role will be responsible for overall stability of the Wealth Management Investment Management application platforms, participation in key optimization initiatives, and collaboration with multiple technical teams within Morgan Stanley. Partner with WM business units, various levels of management and staff to collect, analyze and make recommendations on optimizing the platform. As a team member with expertise in deep analytical triage, you will provide subject matter expertise in debugging, issue analysis and troubleshooting, working with business and technical colleagues to provide reviews and recommendations to avoid any future application issues.

What you'll do in the role:

  • Drive Reliability Engineering Practices
  • Own Production Reliability
  • Production Governance & Change Management
  • Automation & Operational Efficiency
  • Operational Readiness & Knowledge Management
  • Technical Leadership

What you'll bring to the role:

  • 10+ years of experience in a production environment with a solid software development background and understanding of performance tuning, end-to-end troubleshooting, networking fundamentals and appropriate attention to detail
  • BS/MS or equivalent, preferably in quantitative discipline (Computer Science, Computer Engineering).
  • 5+ years' experience in leading a small to medium team of alike skillset.
  • 5+ years of experience in driving SRE principles and Chaos Engineering.
  • Experienced, technically hands-on professional that understands both code and infrastructure
  • Strong experience in scripting language (Shell scripting, Python, Perl, etc.) and cloud driven development
  • Strong database skills with DB2, Sybase or Oracle
  • Hands-on experience with Autosys or other batch scheduling software
  • Experience in AWS/GCP/Azure Cloud technologies
  • Working knowledge on any of the Dev Ops & observability tools (Grafana, Prometheus, Splunk, Kibana)
  • Solid analytical skills, problem determination, and resolution recovery processes
  • Ability to interface and cultivate excellent working relationships with technology teams, business analysts, and vendors
  • Experience in web analytics tools (preferably Adobe Experience Cloud tools) is Plus
  • Should be a fast learner of technologies in a quick paced environment.
  • Have strong organizational skills and the ability to manage multiple tasks and high-pressure situations for outage handling, management, or resolution
  • Is driven to learn about new technologies, techniques and what it takes to be an integral member of this team
  • Hands-on experience administering large-scale, high-availability systems and the tools to monitor performance and availability
  • Excellent communication and writing…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary