Director of Production Support Engineering
Listed on 2026-07-20
-
Software Development
AWS, AI Engineer (Applied/Software), Software Architect
This director role focuses on leading a team of highly skilled production support engineers, defining a long‑term roadmap, ensuring reliability across legacy and modern environments, and driving AI integration.
ResponsibilitiesStrategic Roadmap & Vision:
Devise and champion a long‑term infrastructure strategy that balances rapid AI product innovation with the rigorous demands of enterprise‑grade reliability and global scale.People Leadership:
Manage and grow a world‑class team of Production Support Engineers. Provide technical mentorship, career development, and maintain a high‑performance, low‑ego culture.Multi‑Substrate Scaling:
Lead the transition to a multi‑region and multi‑substrate deployment model within the Salesforce commercial cloud, ensuring resilience across global geographies.Legacy & Custom Operations:
Oversee maintenance and support of legacy environments and bespoke customer deployments, ensuring a seamless experience for long‑term partners during transition to the modern stack.Compliance & Sovereignty Strategy:
Act as the primary stakeholder for Private Cloud Edition (PCE) and Government Cloud initiatives, ensuring infrastructure meets strict regulatory and data‑residency requirements without sacrificing developer velocity.Operational Excellence:
Define and hold the team accountable to SLIs, SLOs, and error budgets. Lead the incident‑to‑insight loop, ensuring production issues result in structural improvements to the product.Stakeholder Management:
Collaborate with Product, Engineering, and Security leadership to align infrastructure investments with the broader business goals of the Agentforce ecosystem.Build and ship high‑quality, production‑grade software using modern engineering practices, integrating AI into the development workflow to deliver secure, optimised, and high‑quality code.
Design and orchestrate complex systems where AI agents integrate seamlessly into human workflows, driving efficiency and innovation at scale.
Contribute to building and maintaining a shared system context—an explicit repository of system designs, constraints, and standards enabling AI to operate accurately and reliably.
Critically evaluate code—human or AI‑generated—for correctness, quality, security, and performance.
10+ years of experience in Production Engineering, SRE, or Infrastructure roles, with at least 3‑5 years in a formal people‑management capacity.
Deep Technical Roots:
Comfortable reviewing architecture designs for Kubernetes, Terraform, and distributed databases such as PostgreSQL, and understanding operational nuances of AI/GPU workloads.Experience at Scale:
Proven track record of managing production environments through a “1 to 100” scaling phase, ideally in a multi‑region or multi‑cloud context.Strategic Clarity:
Ability to articulate a complex technical vision and turn it into an actionable roadmap that stakeholders rally behind.Enterprise Savvy:
Experience supporting large enterprise customers with custom deployment needs and high‑compliance requirements.High‑Stakes Leadership:
Proven ability to lead through major incidents and complex architectural migrations with a calm, decisive approach.AI‑Forward Mindset:
Experience leveraging AI tools to optimise team workflows and a clear vision for how AI will transform the future of production support.Demonstrated AI‑first approach to engineering, using AI to move faster, build fluency across the stack, and contribute beyond specialty.
Experience using AI tools (e.g., Claude Code, Git Hub Copilot, Codex, Cursor, etc.) in development workflows.
Advanced prompt‑engineering skills and the ability to write precise, structured prompts and cultivate system context that makes AI outputs reliable, secure, and production‑ready.
Experience with Government/Sovereign Clouds:
Direct experience navigating FedRAMP, IL5/IL6, or similar high‑security compliance frameworks.Salesforce Infrastructure Knowledge:
Familiarity with Hyperforce, Salesforce Data Cloud, or the internal mechanics of a Private Cloud Edition.Hybrid Infrastructure:
Experience managing the bridge between public cloud (AWS/GCP/Azure)…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).