Manager of Site Reliability Engineering
Listed on 2026-07-21
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Azure
Responsibilities
- Ethoca and Mastercard are seeking a Manager, Site Reliability Engineering, to lead globally distributed application operations teams supporting mission‑critical Java applications across on‑premises infrastructure and Azure public cloud environments
- You will provide leadership, coaching, and technical oversight for teams operating in a 24/7 support model while partnering with Ethoca and Mastercard stakeholders to deliver secure, resilient technology solutions and drive observability, incident management, operational excellence, and continuous improvement
- Lead and manage globally distributed application operations teams supporting mission‑critical Java‑based applications across on‑premises infrastructure and Azure public cloud environments in a 24/7 support model
- Provide strong leadership, coaching, and mentorship to application operations teams, driving employee development and a collaborative, high‑performance culture
- Provide escalation support to the team, including in-depth problem investigation and resolution of complex technical issues
- Establish and promote sustainable incident management practices, including blameless post‑incident reviews, escalation management, and continuous operational and process improvement
- Champion application observability across logging, monitoring, and alerting through APIs and tools such as Azure Monitor, Dynatrace, Splunk, Zabbix, or similar platforms
- Partner with Ethoca and Mastercard stakeholders to design, implement, operate, configure, and optimize secure, resilient, end‑to‑end technology solutions within a highly secure computing environment
- Lead customer and regulatory audits, including PCI‑DSS and SOC 2 compliance and annual disaster recovery exercises, while developing and maintaining high‑quality technical documentation, procedures, and standards
- Manage multiple concurrent projects while ensuring business‑as‑usual operational activities are consistently delivered
The ideal candidate is customer‑ and employee‑focused, intellectually curious, analytical, and comfortable leading through ambiguity in a fast‑paced, highly secure environment. Working knowledge of ITIL, ITSM, and operational best practices, with hands‑on experience using tools such as Remedy, Jira, and Confluence. Highly accountable and outcome‑focused, with the ability to manage multiple initiatives, projects, and team activities in a fast‑paced environment while delivering high‑quality results on time.
Demonstrated technical leadership, coaching, mentorship, delegation, conflict management, and negotiation skills, with excellent written and verbal communication. Strong problem‑solving skills, including the ability to guide complex technical issue resolution, think quickly, and deliver effective solutions under pressure. Proven ability to create, follow, and guide others in adhering to documented processes and procedures while staying current with emerging technologies. Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
Proven experience managing Azure public cloud services, Linux or Unix servers, and Java‑based applications in secure, large‑scale environments. Strong familiarity with observability and monitoring platforms such as Zabbix, Dynatrace, Splunk, Azure Monitor, or similar tools. Strong security and networking knowledge, including encryption, hardening, certificates, vulnerability management, PKI, SSH, VPN, and core networking concepts. Relevant technical certifications in Azure, cloud operations, site reliability engineering, ITIL, security, or related disciplines.
Experience supporting highly secure, regulated, or compliance‑driven environments, including customer or regulatory audits such as PCI‑DSS, SOC 2, or disaster recovery exercises.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: