×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Alpharetta, Fulton County, Georgia, 30239, USA
Listing for: Morgan Stanley
Full Time position
Listed on 2026-06-26
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 110000 - 150000 USD Yearly USD 110000.00 150000.00 YEAR
Job Description & How to Apply Below

Overview

In the Technology division, we leverage innovation to build the connections and capabilities that power our Firm, enabling our clients and colleagues to redefine markets and shape the future of our communities. This Lead Software Production Management & Reliability Engineering position is at Director level and is responsible for overseeing the production environment, ensuring operational reliability of deployed software, and implementing strategies to optimize performance and minimize downtime.

Job Summary

We are looking for a Site Reliability Engineer with a minimum of 5 years of industry experience, preferably in the financial IT community. The role focuses on production support within the WM Product Technology team, automating deployments and working with agile teams to build and support stable and reliable production systems. The ideal candidate will be passionate about automation and proficient in programming languages such as Python, PERL, SHELL, Ruby, Java, or C# and have a strong understanding of database concepts, job schedulers (e.g., Autosys), MQ, web services, UNIX/Linux/Windows OS, and debugging applications.

The candidate should also be a strong leader with excellent communication skills, organized, disciplined, detail‑oriented, self‑motivated, and delivery‑focused.

Responsibilities
  • Maintain applications after deployment by measuring and monitoring availability, latency, and overall system health, focusing on business activities and continuously evaluating cost and TOIL.
  • Engage in and improve the whole lifecycle of services from inception and design, through deployment, operation, capacity planning, and launch reviews.
  • Scale systems sustainably through automation and evolve them by pursuing changes that improve reliability and velocity, including operational automation.
  • Troubleshoot infrastructure issues, review log files, update documentation, and maintain a knowledge base with resolutions.
  • Collaborate closely with the application development team to understand the platform and create tools/utilities that aid production management.
  • Work with upstream data providers and consumers to reduce escalation to development teams.
  • Develop scripts and assist with code changes along with operational tasks and activities.
  • Ensure the support team has excellent knowledge of the application set, owns and maintains the support knowledge base and documents.
  • Use analytical skills to identify trends in the environment and drive problem resolution.
  • Lead efforts to determine improvement areas to stabilize the production environment.
  • Identify risks and act with urgency, working within a team or independently.
  • Test and tune network, hardware, and software configurations to maximize performance.
  • Interface with teams such as IT Dev managers and infrastructure teams and serve as a Subject Matter Expert (SME) for supported applications.
  • Take ownership of and manage production requests, questions, issues, and perform root cause analysis for outages and incidents.
  • Be flexible to provide weekend on‑call rotation and be available for offshore time lead.
  • Be accountable for the Production and non‑Production environments and be part of 24/7 production support coverage.
Skills Required
  • 5+ years of experience in a production environment with a solid software development background and understanding of performance tuning, end‑to‑end troubleshooting, networking fundamentals, and attention to detail.
  • Ability to focus and provide resolutions for production issues in a high‑demand, pressured environment.
  • 5+ years of hands‑on experience designing, developing, and implementing technical solutions, or significant experience in deep technical support.
  • Strong experience in scripting languages (Shell scripting, Python, Perl, etc.) and cloud‐driven development.
  • Strong database skills with DB2, Sybase, or Oracle.
  • Hands‑on experience with Autosys or other batch scheduling software.
  • Strong experience in continuous integration and continuous deployment.
  • Strong experience with on‑demand environments for both virtual machines and containers.
  • Knowledge and hands‑on experience with monitoring tools such as Splunk, IP Soft, Sockeye.
  • Practical…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary