AI Platform Engineering Director; Primarily Office
Listed on 2026-07-23
-
IT/Tech
AI Engineer (Applied/Software), SRE/Site Reliability
Director, Enterprise AI Platform Engineering
This position provides strategic and technical leadership for the enterprise AI platform engineering function, setting the technology strategy and reference architecture for how AI is built, deployed, and scaled enterprise-wide. The role is accountable for building and operating shared AI engineering capabilities that enable teams to develop, deploy, monitor, evaluate, and scale AI solutions safely and efficiently. As Director, you will lead teams responsible for AI platform architecture, reusable engineering frameworks, MLOps/LLMOps, AI operations, agentic workflow patterns, LLM pipeline frameworks, observability, evaluation controls, production support patterns, systems-of-record integration, cost optimization, and technical guardrails that support responsible and scalable AI adoption.
This role works across Technology and business domains, including Information Security, Governance, Enterprise Architecture, Application Development, Digital Services, Infrastructure, Data Engineering, Product, Operations, Risk, Legal, Compliance, and strategic technology partners. You will also manage or coordinates key AI platform partners, including GCP, AWS, Data Dog, Service Now, Salesforce and more where applicable.
Position Compensation Range:
$ - $
Pay Rate Type:
Salary
Compensation may vary based on the job level and your geographic work location. Relocation support is offered for eligible candidates.
Primary Accountabilities- Lead AI platform architecture and strategy
- Define the architecture, standards, and roadmap for shared enterprise AI platform capabilities. This is collaborative with Data & AI Architects.
- Ensure the platform supports scalable, secure, reliable, and governed AI delivery across multiple business domains.
- Balance AI operations, MLOps/LLMOps, and agentic frameworks
- Lead engineering practices for model and AI-system deployment, monitoring, testing, evaluation, versioning, reliability, and lifecycle management.
- Balance operational reliability with reusable agentic workflow patterns and LLM pipeline frameworks.
- Ensure AI systems can be supported and improved after production deployment.
- Contribute to the enterprise AI technology maturity view
- Contribute platform, engineering, operations, observability, support, resilience, cost, systems integration, and production-readiness inputs to the enterprise-level shared AI technology maturity view.
- Use this maturity view to identify capability gaps, guide investment recommendations, and communicate platform and engineering readiness.
- Build reusable engineering frameworks and capabilities
- Build and maintain reusable frameworks, components, and patterns that accelerate AI delivery and reduce duplicated engineering effort.
- Ensure durable reusable capabilities are documented, discoverable, supportable, and governed so they can be leveraged across multiple domains.
- Own platform production support and operational run patterns
- Directly own production support for the AI platform and shared AI engineering capabilities, especially L1, L2, and the engineering side of L3 support.
- Establish production run patterns for AI-enabled workflows, including support models, incident paths, escalation patterns, fallback mechanisms, and human handoff design.
- Partner with customer service, employee support, application support, service management, or other operational channels where AI experiences require context-rich support transitions.
- Establish observability, evaluation, and monitoring capabilities
- Establish platform capabilities for telemetry, model and system monitoring, traceability, drift or quality signals, evaluation controls, regression checks, and operational visibility.
- Provide visibility into AI system behavior, performance, usage, and risk signals.
- Embed responsible AI technical guardrails
- Embed responsible AI technical controls into platform capabilities in alignment with shared enterprise governance expectations.
- Ensure platform capabilities support auditability, policy adherence, access controls, least-privilege operation, model/agent registration, and responsible AI practices.
- Lead AI cost, capacity, and resilience practices
- Provide…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).