×
Register Here to Apply for Jobs or Post Jobs. X

Sr. Manager, Site Reliability Engineering

Job in Coppell, Dallas County, Texas, 75019, USA
Listing for: Brinker International
Full Time position
Listed on 2026-06-14
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Project Manager
Salary/Wage Range or Industry Benchmark: 100000 - 125000 USD Yearly USD 100000.00 125000.00 YEAR
Job Description & How to Apply Below

Sr. Manager, Site Reliability Engineering

Coppell, TX

Job Summary

We are seeking a highly skilled and motivated Sr. Manager, Site Reliability Engineering to build and lead our Site Reliability Engineering capability from the ground up. In this role, you will contribute to the vision and define operating model and foundational practices needed to improve the reliability, scalability, and performance of our technology platforms and services. You will establish core SRE disciplines such as automation, observability, incident response, service level objectives, and capacity planning while partnering closely with infrastructure, development, and operations teams to embed reliability engineering across the software lifecycle.

This leader will be responsible for building a high-performing team, implementing scalable processes and tooling, and driving continuous improvement that reduces operational toil, strengthens resilience, and supports the evolving needs of the business.

This role is based in Dallas (Coppell), TX and follows a hybrid schedule (3 days in office). We are currently focused on local candidates or those open to relocating to the area at their own expense. At this time, we are unable to provide sponsorship support.

Objectives

  • Build and lead the Site Reliability Engineering capability from the ground up by establishing the team, operating model, standards, tooling, and foundational processes needed to support scalable, reliable platform infrastructure and applications.
  • Drive reliability, availability, and delivery performance by implementing automation, observability, incident response practices, and service level objectives in partnership with infrastructure, development, and operations teams.
  • Continuously improve system performance, resilience, and operational efficiency through proactive monitoring, root cause analysis, capacity planning, and data-driven optimization that reduces toil and supports evolving business needs.

Your Key Job Functions

  • Build, lead, and mentor a high-performing Site Reliability Engineering team, establishing clear priorities, accountability, and engineering standards to support a scalable and resilient operating model.
  • Define and implement the foundational SRE strategy, including service level objectives, reliability requirements, operating processes, and governance in partnership with infrastructure, development, and operations teams.
  • Design, implement, and maintain scalable and reliable infrastructure to support our applications and services.
  • Develop and maintain automation for deployment, monitoring, incident response, and operational workflows to reduce toil and improve consistency, speed, and reliability.
  • Lead or provide input into incident response and problem management practices, including root cause analysis, corrective actions, and prevention strategies to improve service availability and resilience.
  • Establish and optimize observability practices by gathering and analyzing metrics, logs, and system telemetry to support performance tuning, fault isolation, and proactive issue detection.
  • Partner with development and IT teams to embed reliability, testing, release discipline, and operational readiness into the software development lifecycle.
  • Gather and analyze metrics from operating systems, logs, as well as applications to assist in performance tuning and fault finding.
  • Partner with development teams to improve services through rigorous testing and release procedures.
  • Drive continuous improvement in system performance, scalability, and capacity through monitoring, testing, optimization, and data-driven operational insights.
  • Provide technical leadership in system design, platform management, and capacity planning while balancing feature delivery speed with reliability and service commitments.

What You Bring to the Team

  • Master's degree and/or bachelor's degree in combination with equivalent experience in Computer Science, Engineering, or related field.
  • 5+ years as a Site Reliability Engineer or similar role, with a demonstrated track record of successfully managing reliability and scalability of large-scale systems.
  • Strong knowledge of cloud platforms (
    AWS
    , Azure
    , Google Cloud
    ) and containerization technologies (
    Docker
    , Kubernetes
    ).
  • Proficiency in scripting and automation languages (
    Python
    , Bash
    , Ansible
    ).
  • Experience with monitoring and logging tools (
    New Relic
    , Data Dog
    , Prometheus
    , Grafana
    , ELK stack
    ).
  • Demonstrated leadership experience, with a passion for mentoring and developing team members.
  • Excellent problem-solving skills and the ability to work under pressure.
  • Proven ability to solve complex issues in a timely fashion.
  • Proven ability to quickly adapt and flex to a dynamic environment by being a "self-starter".
  • Strong communication and collaboration skills.
  • Strong project management skills.
  • Strong documentation skills.
  • Solid understanding of networking, security, and system administration.
  • Experience with infrastructure as code (IaC) tools (
    Terraform
    , Cloud Formation
    ).
  • Knowledge of CI/CD…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary