×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineering; SRE) Manager

Job in Frisco, Collin County, Texas, 75034, USA
Listing for: McAfee GmbH
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Project Manager
Salary/Wage Range or Industry Benchmark: 123650 - 229650 USD Yearly USD 123650.00 229650.00 YEAR
Job Description & How to Apply Below

Role Summary

It's an exciting time to join McAfee!

We're looking for an experienced SRE Manager to lead our growing North American Site Reliability Engineering team and own reliability strategy across our Cloud and Kubernetes platforms. You'll combine deep technical credibility with strong people leadership — setting direction, growing the team, and acting as the senior escalation point during the most critical incidents — while partnering closely with engineering and business leadership.

This is an onsite position located in our Frisco, TX office. We are only considering candidates within a commutable distance to the Frisco office.

Position Details

About the role:
  • Lead and grow a team of SREs, setting technical direction and reliability strategy across AWS, GCP infrastructure and EKS platforms, and GKE platforms.
  • Own the organization's Incident and Problem Management processes, ensuring major incidents are handled efficiently, with timely executive communication and thorough post-incident reviews.
  • Drive the team's automation strategy, championing Python-based tooling and frameworks that reduce manual toil and improve reliability at scale.
  • Set standards for Terraform-based infrastructure-as-code, ensuring secure, scalable, and consistent provisioning practices across teams.
  • Define the organization's observability strategy, ensuring Grafana dashboards, alerting, and SQL/Cloud Watch-based analysis practices scale effectively.
  • Act as a senior escalation point and incident commander for the most critical, high-severity incidents, providing calm, decisive leadership under pressure.
  • Partner with senior leadership, product, and engineering stakeholders to communicate risk, reliability posture, and remediation roadmaps clearly and confidently.
  • Own hiring, mentoring, performance management, and career development for the SRE team.
  • Design and build self-healing automation and runbooks that detect known failure patterns and trigger remediation automatically, reducing manual intervention and recovery time for recurring incidents.
  • Implement and maintain monitoring across multiple regions to ensure consistent visibility into system health, latency, and failover readiness across all deployment zones.
  • Proactively identify potential failure points and performance bottlenecks before they impact production and reduce operational workload by automating recurring manual tasks.
  • Drive ITSM process maturity across the organization, partnering with other teams to embed Incident and Problem Management best practices.
  • Manage on-call structure, escalation paths, staffing, and operational readiness for the team.
  • Report on reliability metrics, incident trends, and improvement initiatives to senior leadership.
About You:
  • 9+ years of experience in Site Reliability Engineering, Dev Ops, Infrastructure, or related roles, including significant experience in a leadership or management capacity.
  • Demonstrated experience building, leading, and growing high-performing technical teams.
  • Strategic reliability leadership:
    Balances operational excellence with long-term reliability improvements, ensuring the team addresses immediate risks while building scalable, sustainable practices.
  • Deep, hands-on background with AWS infrastructure and strong technical credibility to guide architecture and operational decisions.
  • Proven track record leading teams through complex EKS/GKE troubleshooting and operational challenges.
  • Strong technical fluency in Python for automation and Terraform for infrastructure-as-code, with the ability to guide and review the team's work.
  • Extensive experience owning ITSM processes — Incident and Problem Management — at an organizational level.
  • Strong command of observability practices, including Grafana dashboards, SQL, Cloud Watch…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary