Head of AI & Agentic Platform Engineering
Listed on 2026-07-18
-
Software Development
AI Engineer (Applied/Software), DevOps, Machine Learning/ ML Engineer
Role Summary
The Head of AI & Agentic Platform Engineering owns the infrastructure layer that makes Pfizer's AI ambitions executable, the compute, LLM gateway, MLOps machinery, and observability platform on which every AI workload at Pfizer runs. This is not a supporting function – it is the capability that determines whether Pfizer's AI strategy moves at the speed of ambition or the speed of infrastructure constraints.
The platform this team builds is the difference between a data scientist who spends two weeks provisioning an environment and one who is running experiments on day one, and between an AI model that takes six months to reach production and one that ships in days through a governed, automated deployment pipeline. The scope of AI workloads this platform must support is broad.
Each Pfizer domain (i.e., R&D, Commercial, Global Supply, Enabling Functions) has distinct compute, latency, governance, and reliability requirements, and this platform must serve all of them without compromise. As Pfizer advances from assistive AI tools toward autonomous agentic systems that take multi-step actions across the enterprise, the demands on this platform will grow in both complexity and consequence. The LLM gateway, agent orchestration layer, and observability infrastructure this leader builds today must be architected for that future from the outset.
- Gateway & Serving:
Architect and maintain the Enterprise LLM gateway with access control, multi-model routing, rate limiting, cost attribution, and audit logging for all LLM interactions across Pfizer, including agentic AI workloads. Provide model serving infrastructure with low-latency inference, auto-scaling, and multi-region deployment for production models. - Agentic AI Runtime:
Build the infrastructure layer that supports autonomous AI agents taking multi-step actions across Pfizer's systems. Enable stateful process management, short-term and long-term memory, tool-calling orchestration, and coordination with other agents. - Observability:
Deliver real-time usage monitoring, cost attribution by team and use case, and anomaly detection for both gateway and agentic operations. - Enterprise Tool & MCP Registry:
Govern the catalog of tools, APIs, and data sources that AI agents are permitted to call at runtime, ensuring compliance with Trusted AI governance. - Compute & Environments:
Provision and manage enterprise compute infrastructure (GPU, TPU, CPU) across cloud and on-premises, including capacity planning, Fin Ops governance, and utilization optimization. Provide pre-configured AI environments and automated IaC provisioning across dev, staging, and prod. - Runtime Enablement:
Deliver an MLOps platform with experiment tracking, model versioning, automated evaluation, deployment pipelines, and model registry, integrating with Trusted AI risk classification and sign-off. - Registry, Deploy & Trust:
Own the Enterprise AI model registry, document every AI model and agent across lifecycle, enforce deployment pipelines with Trusted AI gates, and implement guardrails and policy enforcement for AI governance.
- 12+ years in software or infrastructure engineering, with 7+ years in AI/ML platform, MLOps, or AI infrastructure roles at significant scale.
- Demonstrated experience building and operating multi-tenant AI/ML platform infrastructure, compute provisioning, model training pipelines, model serving, and production monitoring.
- Deep hands‑on experience with LLM gateway or model serving infrastructure, multi-model routing, inference optimization, access control, and cost attribution at enterprise scale.
- Proven MLOps platform experience with documented outcomes in deployment velocity, reliability, and developer satisfaction.
- Strong IaC practices in a multi-cloud architecture (Azure, AWS, GCP) including Terraform expertise.
- Experience leading platform teams with an SLA-driven, product-minded operating model.
- Demonstrated ability to collaborate across organizational boundaries, with adjacent platform teams, security functions, and governance stakeholders.
- Experience operating AI infrastructure in a regulated environment with GxP controls, audit trail requirements, and validated environment obligations.
- Broad leadership experience, including influencing peers, developing and coaching teams, and achieving meaningful business impact.
- Experience building or operating ML platform infrastructure at a major technology company at petabyte scale with thousands of concurrent ML engineers.
- Experience designing agentic AI infrastructure – orchestration layer, memory architecture (short-term, long-term), tool-calling and MCP integration, agent-to-agent communication, and safety architecture.
- Deep LLM-specific infrastructure experience: KV cache management, speculative decoding, quantization trade-offs, and concurrent multi-model serving.
- HPC environment experience, job schedulers (SLURM, LSF, or equivalent), parallel file systems, and large-scale…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).