Lead Support Engineer
Listed on 2026-10-08
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support
WPP is the trusted growth partner for the world’s leading brands. With exceptional talent, trusted data and intelligence, and world-class partnerships – all united by our pioneer agentic marketing platform, WPP Open – we help clients navigate change, capture opportunity, and deliver transformational growth.
WPP Media is WPP’s AI-driven media operating unit, bringing together media, data, and partnerships to deliver creative personalisation nected through WPP Open and powered by Open Intelligence, clients see exactly where, how, and why their media investment is working.
This role is within WPP Media, where you will be instrumental in owning production stability, observability, and system health for a mission-critical global campaign governance and compliance platform. As Lead Support Engineer, you will provide advanced L2+ and L3-oriented support for a system running natively on Google Cloud Platform (GCP). You will investigate complex production issues, implement minor code-level fixes in Python, lead root cause analysis, and improve the reliability of distributed systems.
Working closely with the EMEA-based Engineering Tech Lead, Product Owner, and Quality Assurance partners, you will connect production insights with technical roadmap priorities.
The opportunity combines hands-on Site Reliability Engineering (SRE) with technical squad leadership. A foundational monitoring and alerting setup is already in place, giving you the platform to evaluate the current approach, define a clear observability direction, and strengthen system health monitoring across the environment. You will automate runbooks, reduce manual operational effort, and help shape the future support squad for a dedicated product used across global advertising campaigns.
Key Responsibilities- Own production stability, system health, and observability for a mission-critical campaign governance and compliance platform running on GCP.
- Diagnose and resolve complex, intermittent, and high-priority incidents across application, database, infrastructure, and networking layers.
- Read, debug, and implement minor fixes and patches directly within the existing Python production codebase.
- Define and advance the observability strategy using GCP Cloud Logging, Cloud Monitoring, Error Reporting, Prometheus, PromQL, and suitable service level indicators and objectives.
- Lead incident response, post-incident reviews, and end-to-end root cause analysis, partnering with the EMEA Engineering team on permanent remediation.
- Monitor execution flows, performance, data refreshes, automated checks, and recurring compliance reporting to ensure reliable system operation.
- Identify opportunities to automate manual support activity, develop and maintain runbooks, and reduce operational toil.
- Provide technical leadership to the L2+ support squad through coaching, knowledge sharing, prioritization, and effective handovers.
- Partner with Product, Engineering, and Quality Assurance teams to assess operational risk, improve release and change management, and influence technical roadmap decisions.
- Advanced education in computer science, software engineering, information technology, or a related technical discipline, or equivalent practical experience.
- Strong, current Python proficiency, including the ability to read, debug, and implement fixes directly within a shared production codebase.
- Deep, hands-on experience supporting enterprise systems on Google Cloud Platform, including Compute Engine, Google Kubernetes Engine, Cloud SQL, Big Query, Pub/Sub, Firestore, Cloud Functions, Identity and Access Management, and virtual private cloud networking.
- Complex production troubleshooting and root cause analysis experience across distributed…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).