Senior Manager, Product Software Engineering - Foundations Engineering
Listed on 2026-09-06
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Summary
Sr. Manager, Foundations Engineering leads the day-to-day engineering execution, operational reliability, and platform delivery of the Foundations Engineering function within Ad Platforms - the internal capability layer that powers every service, application, and system in the organization. The organization owns and operates portions of the Ad Platforms infrastructure and developer platform stack directly, while also extending, operationalizing, and supporting shared enterprise platforms within the Ad Platforms ecosystem.
This includes leadership across platform engineering and Dev Ops, cloud infrastructure and runtime operations, SRE and reliability practices, database administration, observability enablement, and the infrastructure and platform capabilities supporting the Data Platform, BI ecosystem, and AI agentic and real-time ML platform. This role partners closely with centralized Platform Engineering and Infrastructure teams responsible for shared enterprise services and cloud platforms.
This role reports to the Director, Product Software Engineering - AI Engineering + Foundations Engineering. We are seeking a Sr. Manager who is deeply execution-oriented and operationally excellent - someone who thrives on building reliable, scalable systems, leading high-performing engineering teams through delivery, and improving platform quality and developer experience quarter over quarter. If you are a leader who takes pride in production stability, drives operational discipline without creating bureaucracy, and knows how to develop engineers while keeping the platforms running at scale, this is an exceptional opportunity.
and Duties of the Role:
Platform Engineering & Dev Ops Manages and delivers the platform engineering and Dev Ops capabilities that power software delivery across Ad Platforms - including ownership and operational support of CI/CD systems, Kubernetes and container orchestration environments, Infrastructure-as-Code modules, deployment automation, shared cloud platform integrations across AWS (primary), Azure, and GCP, and internal developer tooling - improving deployment speed, reliability, scalability, and developer experience.
Drives automation-first operational practices, golden-path engineering patterns, self-service engineering capabilities, and Ads-specific platform abstractions built on top of shared infrastructure platforms - accelerating engineering velocity, reducing operational toil, and improving onboarding, deployment reliability, and release consistency across Ad Platforms teams. Serves as a primary operational ownership layer for Ad Platforms workloads running on shared infrastructure and platform services - partnering closely with centralized Platform Engineering organizations responsible for core cloud, Kubernetes, CI/CD, and observability platforms.
Implements and enforces secure-by-default platform standards across Ad Platforms environments and delivery pipelines - including secrets management, deployment governance, engineering quality gates, vulnerability remediation, auditability, and compliance controls - in partnership with Security and centralized infrastructure teams. Develops internal developer tooling and web-based operational interfaces that support service deployment, runtime management, and engineering productivity across Ad Platforms.
Reliability, Data & AI Platform Infrastructure Establishes and operates reliability engineering practices across Ad Platforms foundational services and runtime environments - including SLO/SLI frameworks, incident management rigor, production readiness standards, operational health monitoring, and resiliency engineering practices - driving measurable improvements in uptime, stability, and deployment confidence. Leads observability operations and enablement across Ad Platforms - including logging, monitoring, distributed tracing, alerting frameworks, and operational instrumentation standards - partnering with centralized observability teams to ensure consistent and actionable visibility into system behavior and platform health.
Leads database engineering operations and data infrastructure enablement for Ad Platforms, including operational reliability, backup and recovery strategy, failover readiness, performance optimization, and modernization efforts across relational, No
SQL, and cloud-native data systems supporting transactional services, analytics, BI, and Data Platform workloads. Builds and operates infrastructure and platform capabilities supporting AI agentic and real-time ML systems - including orchestration runtimes, inference platform enablement, GPU and compute orchestration, deployment automation, and governance-aligned operational frameworks - enabling AI Engineering teams to develop, deploy, and scale AI systems reliably and safely. Executes Fin Ops and cost management practices across multi-cloud infrastructure and AI inference workloads - including capacity planning, cloud spend visibility,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).