Principal AI Platform Engineer
Listed on 2026-06-26
-
IT/Tech
AI Engineer (Applied/Software)
Why We Need You
We are seeking a Principal AI Platform Engineer to join our community. As AI agents become mission‑critical in production business workflows, PEMCO needs a leader who owns the operational reliability, governance, security, and cost management of the AI layer. In year one, this is a hands‑on technical leadership role. You will build systems yourself while establishing the standards and governance framework.
As AI operations mature, you will build and scale a team to match operational demands.
The role spans both business AI use cases (in partnership with the Data, AI & Digital teams) and technology enablement use cases including IT operations, information security, help desk automation.
You will be responsible for building and maintaining the enterprise agent marketplace, establishing production‑grade observability, and ensuring governance and compliance across all AI deployments.
What You’ll Be Doing- Own the operational lifecycle of AI agents deployed across PEMCO: deployment, monitoring, scaling, incident response, and retirement.
- Build and maintain observability for the AI layer, including cost tracking, latency, error rates, model performance, token usage, and production monitoring for agent workers.
- Manage agent orchestration infrastructure, including configuration, versioning, connection management, and tool registration. Current stack includes MCP‑based orchestration and Azure OpenAI services.
- Establish runbooks and incident response procedures for AI agent failures. When an agent supporting a business workflow goes down, this role owns the recovery.
- Implement prompt governance controls and role‑based model access per Information Security standards: PII exposure monitoring, prompt injection detection, and access enforcement. Contribute requirements and technical capabilities to Info Sec for AI‑specific policy development.
- Build the enterprise agent marketplace: deploy and manage AI agents within an enterprise UI framework, ensuring discoverability, versioning, and access controls.
- Evaluate new model releases, track capability evolution, and make recommendations on model selection. Maintain a knowledge management layer for AI operations including decision logs, model inventories, and governance documentation.
- Contribute directly to AI governance and the AI Governance Working Group. Operationalize security, governance, and compliance standards defined by Information Security and Data & AI leadership across all AI deployments.
- Optimize AI infrastructure costs through model selection, caching strategies, batching, and token budget management.
- Partner with the Data & AI team on production readiness for models and agents. Data & AI owns model development, training, and AI governance policy. This role owns the operational deployment, monitoring, fallback behavior, graceful degradation, and knowledge layer integration.
- As the function matures, define team structure, secure headcount, and build a team.
- Demonstrate behaviors consistent with PEMCO’s policies, values, code of ethics, and business conduct.
- Authentically support the PEMCO Brand and constantly be on the lookout for top talent to join us to achieve our Mission to Worry Less and Live More.
- Other duties as assigned.
- Technical degree or equivalent practical experience.
- 5+ years in a technical operations, platform engineering, or SRE leadership role is required.
- 2+ years building and deploying AI agents in production environments, including AI/ML ops (model monitoring, feature store management, RAG, vector store enablement) is required.
- Experience with cloud AI services (Azure OpenAI, AWS Bedrock, Google Vertex AI, or comparable) is required.
- Experience building observability and monitoring for production services is required.
- Experience with identity and access management for technical platforms is required.
- Demonstrated understanding of AI security risks: prompt injection, data leakage, model abuse is required.
- Experience leading or building a technical team is required.
- Experience with agent orchestration frameworks (MCP, Lang Chain, CrewAI, Auto Gen, or similar).
- Experience with LLM operational patterns: token management,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).