Director, Platform Engineering
Listed on 2026-07-02
-
IT/Tech
SRE/Site Reliability
Ingram Micro is seeking a Director, Platform Engineering to own both the day-to-day operational health of our technology platform and the long-term transformation of how we build, release, and run software. This is a dual-mandate leadership role: keep the lights on and evolve what those lights look like.
This role sits within the Global Platform Technology organization, part of the broader CTO organization. The Director will report directly into the Executive Director, Agile Delivery & Program Management, who leads the full cross‑functional technology enablement portfolio alongside Engineering. This structure ensures the role is tightly connected to both engineering execution and the broader operational and delivery agenda.
The Director will lead four core disciplines; SRE, Dev Ops, Performance Engineering, and Release Management, with a clear, unifying transformation agenda: replace manual processes with intelligent automation, embed AI across every operational function, and build a technology operations organization that is self‑sufficient, preventive, and scalable. The answer to operational complexity is not more people, it is smarter systems.
Your role : THE DUAL MANDATE- RUN:
Day-to-Day Operations
- Own the operational stability, reliability, and performance of the platform. Ensure SRE, Dev Ops, release, and performance functions deliver consistently against SLOs, release cadences, and quality benchmarks. - TRANSFORM: AI & Automation Agenda
- Lead the transformation of all four disciplines into AI‑powered, self‑serve functions—where automation handles the routine, AI anticipates the unexpected, and engineers focus on outcomes rather than operations.
Site Reliability Engineering (SRE)
- Implement AI‑powered predictive monitoring and anomaly detection to surface degradation signals before they become incidents.
- Build automated remediation, self‑healing infrastructure, AI‑triggered rollbacks, and automated runbook execution, to resolve issues without human intervention.
- Shift the SRE model from reactive firefighting to proactive reliability engineering, using AI insights to eliminate recurring failure patterns.
- Govern SLIs, SLOs, and error budgets as the framework for reliability investment, with real‑time AI forecasting of compliance.
- Embed reliability thinking early in the engineering lifecycle—shifting reliability left, not bolting it on at the end.
- Design and execute an AI‑driven CI/CD strategy targeting 100% touchless, end‑to‑end pipeline automation, zero manual intervention from commit to production.
- Embed AI across the pipeline for intelligent code analysis, automated testing, self‑healing deployments, and anomaly detection.
- Architect and deliver self‑serve developer platforms that empower teams to provision environments, trigger deployments, and manage releases independently.
- Define milestones and metrics for the automation roadmap, tracking progress toward full pipeline autonomy.
- Champion a culture of engineering self‑sufficiency, reducing toil and enabling teams to ship faster with greater confidence.
- Own load, stress, and scalability testing programs, ensuring the platform performs reliably under current and projected demand.
- Automate performance testing as a continuous activity within CI/CD pipelines, blocking releases that regress against defined baselines.
- Use AI to identify performance bottlenecks, forecast capacity needs, and recommend optimizations before issues surface in production.
- Drive Fin Ops practices to optimize cloud resource utilization and infrastructure spend alongside performance improvements.
- Own the end‑to‑end release process across environments, ensuring consistency, auditability, and minimal disruption to production.
- Automate release orchestration, approval workflows, and deployment coordination to reduce manual overhead and human error.
- Implement AI‑assisted release risk scoring to evaluate change impact, recommend optimal deployment windows, and flag high‑risk releases before they go out.
- Establish feedback loops between release outcomes and development practices to continuously reduce change failure…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).