×
Register Here to Apply for Jobs or Post Jobs. X

Software Engineering Manager – Site Reliability Center

Job in Auburn, Lee County, Alabama, 36831, USA
Listing for: Jobtailor
Full Time position
Listed on 2026-07-20
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 120000 - 160000 USD Yearly USD 120000.00 160000.00 YEAR
Job Description & How to Apply Below

Responsibilities

  • Manage SRE and related teams; lead, coach, and develop a team of SRE engineers; set clear goals, drive accountability, and foster a culture of ownership and excellence; partner with cross‑functional stakeholders to align technology and business objectives; support talent development, performance management, and succession planning; encourage innovation, continuous learning, and Dev Ops/SRE best practices.
  • Lead incident management & remediation; manage and actively participate in end‑to‑end incident response for major (P1/P2) incidents; guide real‑time triage, diagnostics, and troubleshooting across application, infrastructure, and network layers; ensure rapid execution of remediation actions and service restoration; provide clear, timely communication to stakeholders during incidents; oversee post‑incident analysis, reporting, and documentation to drive improvements.
  • Provide technical leadership in production support; serve as an escalation point for complex production issues; guide troubleshooting across applications, infrastructure (Linux/Windows), databases (Oracle, SQL), middleware and integrations; ensure efficient log, metric, and system analysis; oversee batch/ETL monitoring and recovery processes; foster strong collaboration across engineering, infrastructure, and vendor teams.
  • Drive problem management & root‑cause resolution; lead root‑cause analysis (RCA) efforts for major and recurring incidents; ensure ownership and resolution of problem records; drive permanent fixes and systemic improvements to eliminate repeat issues, identify trends and patterns to reduce risk and improve stability; partner with engineering teams to resolve code defects and system gaps and promote knowledge sharing via runbooks, knowledge articles, and error catalogs.
  • Oversee change management & release execution; ensure safe and compliant execution of production changes and releases; validate change readiness, testing, rollback strategies, and risk assessments; represent the team in CAB reviews, providing technical risk evaluation; oversee post‑implementation reviews (CPIR) and ensure follow‑through and drive improvements in change success rate and reduction in production defects.
  • Advance monitoring, alerting & observability; lead efforts to build and optimize monitoring, dashboards, and alerting frameworks, champion use of tools such as Dynatrace, Big Panda, Logscale, and enterprise platforms, improve signal‑to‑noise ratio through alert tuning; enable proactive issue detection before customer impact; strengthen event management and observability practices.
  • Champion resiliency, stability & availability; lead efforts to ensure high availability of critical systems; oversee disaster recovery, failover, and continuity testing; identify and eliminate single points of failure and drive improvements in MTTR, uptime, and service reliability.
  • Enable scalability & performance optimization; guide capacity planning and performance tuning strategies; ensure systems scale effectively under peak demand; partner with development teams for performance‑driven design improvements; optimize system configurations to improve efficiency and throughput.
  • Lead a 24x7 production support model; manage team participation in a 24x7 on‑call rotation; oversee engagement in incident bridges, war rooms, and escalations; support pod‑based operating models aligned to key applications; ensure seamless handoffs and global support continuity.
  • Drive Automation & Operational Efficiency; identify and prioritize opportunities to reduce manual effort through automation; implement automation across incident remediation, monitoring and alerting, deployment and validation, promote standardized runbooks and automation frameworks and improve operational metrics and reduce toil.
  • Ensure Governance, Risk & Compliance; maintain adherence to enterprise policies and regulatory standards; support audits, vulnerability remediation, and risk controls; ensure accurate documentation and operational procedures and champion security, access management, and data governance practices.
Requirements
  • 5 + years of related experience and 3+ years of management experience.
  • Strong experience in Site Reliability Engineering, Production Support, or Dev Ops.
  • Proven ability to lead teams in high‑availability, enterprise environments.
  • Deep understanding of incident, problem, and change management frameworks.
  • Hands‑on knowledge of monitoring tools, cloud/infrastructure platforms, and automation.
  • Experience improving system reliability, observability, and operational maturity.
  • Strong communication skills with the ability to lead during high‑pressure situations.
  • Experience with OCP under infrastructure (Linux/Windows, OCP), MongoDB, Cassandra under databases (Oracle, SQL, MongoDB, Cassandra) and working knowledge of Elasticsearch, Redis, MQ and Kafka is a plus.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary