Senior Director, Enterprise Reliability Engineering
Listed on 2026-07-24
-
IT/Tech
Systems Engineer, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Senior Director, Enterprise Reliability Engineering
The Senior Director, Enterprise Reliability Engineering is responsible for improving the reliability, resiliency, scalability, and operational excellence of Marriott's Data and Core technology domains. This role plays a critical part in Marriott's reliability transformation, shifting the organization from reactive, centralized operations toward engineering-owned reliability grounded in modern Site Reliability Engineering (SRE) and Dev Sec Ops practices. This leader serves as a catalyst for change, partnering closely with product, engineering, platform, architecture, and security teams to embed reliability throughout the software development lifecycle.
Success is measured by measurable improvements in reliability outcomes, incident reduction, engineering adoption, and operational maturity rather than ownership of ticket queues or production support.
Candidate Profile
Required:
- Undergraduate degree in Engineering, Information Systems or Computer Science discipline and/or equivalent experience/certification
- 10+ years of senior IT leadership experience across software engineering, platform engineering, SRE, Dev Ops, or related disciplines with a blend of deep technical knowledge and a customer-focused mindset
- 8+ years leading large-scale engineering or reliability organizations
- 3-5 years' experience operating and maintaining a Multi-Cloud environment
- 8+ years managing direct reports and Service Provider teams
- Demonstrated success driving reliability transformation across multiple engineering domains
- Deep understanding of cloud-native architectures, distributed systems, and modern software delivery
- Proven ability to influence technical strategy and standards without direct authority
- Experience defining SLOs, SLIs, error budgets, and operational maturity frameworks
- Strong executive communication skills and systems thinking mindset
Preferred:
- Graduate Degree in a technical discipline
- 3+ years experience managing Public Cloud technology stacks such as AWS, Azure, Alibaba, GCP, etc.
- 10+ years experience leading and engaging highly technical architecture, engineering and operations teams
- 10+ years of hands on technical experience in infrastructure and/or development teams
- Demonstrated ability to manage a large, diverse portfolio of technologies and projects, while balancing short-term goals and long-term vision
- Comfortable applying a combination of qualitative and quantitative methods to define success
- Expert problem solver that is to be able to solve challenges across people, process and technology
- Ability to build strong relationships and network throughout the broader IT community
- Experience in the security, implementation and operational support of mission critical products
- Strong influencing skills and an ability to overcome barriers while driving change
- Experience in researching emerging technologies and trends, standards, and products
- Experience in developing technology roadmaps and strategies
- Excellent verbal and written communication skills for a wide range of audiences including executives, business stakeholders and IT teams
- Strong knowledge of emerging tools, software, applications, and systems for attaining best-in-class IT technology across the enterprise
- Strong attention to detail with an ability to operate effectively across multiple priorities
- ITILv4 certification
- Experience operating in an Agile or SAFE framework
- AWS, Azure, Oracle or similar cloud certifications
- Familiarity with hospitality, travel, retail, or large-scale consumer-facing platforms
Core Work Activities
Domain Reliability Strategy & Transformation
- Define and execute domain reliability strategies for Data, and Core platforms
- Translate enterprise reliability transformation objectives into actionable domain roadmaps
- Establish service ownership models, service tiering, and reliability targets
- Identify organizational, process, and technology barriers to reliability adoption and drive remediation
SRE & Dev Sec Ops Enablement
- Expand and mature embedded SRE capabilities within domain engineering teams
- Drive adoption of SLOs, error budgets, production readiness, and operational excellence practices
- Embed reliability considerations into architecture reviews and CI/CD pipelines
- Improve deployment safety, change quality, and risk reduction
Incident Reduction, Resiliency & Learning
- Lead blameless post-incident reviews and root cause analysis
- Ensure systemic corrective actions that prevent repeat incidents
- Promote resiliency engineering practices such as fault isolation and graceful degradation
- Continuously improve incident response effectiveness and prevention
Enterprise Engineering Excellence & Automation
- Promote resiliency engineering practices such as fault isolation and graceful degradation
- Eliminate operational toil through automation and self-healing systems
- Establish and prioritize domain reliability backlogs focused on systemic risk reduction
- Increase engineering leverage through reusable reliability patterns and platforms
- Advance…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).