Senior Manager, AI Reliability Engineering -Kroger Technology & Digital (P2498
Listed on 2026-08-21
-
IT/Tech
SRE/Site Reliability, AI Engineer (Applied/Software), AI Business & Operations
Senior Manager, AI Reliability Engineering
The Senior Manager, AI Reliability Engineering at Kroger Technology & Digital (KTD) leads the engineering discipline that makes enterprise AI operationally trustworthy the enterprise moves from AI pilots to production systems that associates and customers depend on every day, this leader ensures that models, agents, copilots, AI gateways and shared runtime services are dependable, high-quality, responsive and cost-efficient by design. This leader will stand up a new discipline from the ground up: defining what production-grade AI means, engineering the standards and automation that bake reliability and quality into every system, and shaping how build teams design for resilience from day one.
The role blends engineering leadership, technical ownership and cross-functional influence. Reliability is a strategic enabler of adoption and velocity - the difference between experimenting with AI and confidently running it at scale.
Key responsibilities include building the discipline (0-to-1), defining what production-grade, operationally trustworthy AI means for the enterprise, standing up the AI Reliability Engineering function, positioning reliability as an enabler of AI adoption and velocity, engineering reliability and quality into systems, providing production readiness and agent onboarding, owning production quality, observability and agentops, driving efficiency and performance, providing resilience and incident excellence, and leading and influencing.
Required qualifications include minimum 10+ years of experience in software, platform, cloud, SRE or production engineering, including 3 or more years leading engineering teams. Strong technical depth in distributed systems, Kubernetes, cloud infrastructure on GCP and/or Azure, networking, identity, automation, CI/CD and modern observability engineering is also required. Experience establishing SLOs, error budgets, on-call practices, production-readiness reviews, incident management and post-incident improvement programs is necessary.
Direct experience with ML, AI or LLM systems in production, or demonstrated ability to master failure modes unique to models and agents is required. Ability to influence architecture and roadmaps, lead through ambiguity, partner across a matrixed organization and drive outcomes without direct authority is also required. Excellent technical and executive communication is needed, with the ability to shape strategy with leaders and go deep with senior engineers.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).