Cloud Operating Model - Managing Consultant
Listed on 2026-09-01
-
IT/Tech
SRE/Site Reliability, AI Business & Operations, AI Engineer (Applied/Software)
ABOUT CAPGEMINI
At Capgemini Invent, we believe difference drives change. As inventive transformation consultants, we blend our strategic, creative and scientific capabilities, collaborating closely with clients to deliver cutting‑edge solutions. Join us to drive transformation tailored to our client's challenges of today and tomorrow. Informed and validated by science and data. Superpowered by creativity and design. All underpinned by technology created with purpose.
YOURROLE
As an AI Platform & Site Reliability Engineering Managing Consultant, you will help clients design, build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production‑grade, enterprise‑scale AI services.
You will work with technology, engineering, operations and business leaders to establish the platforms, operating models, governance and reliability practices required to run AI‑enabled services safely, effectively and ing as a trusted advisor to senior stakeholders, you will shape client strategy while leading delivery teams and helping grow our AI Platform & Reliability Engineering capability.
This will include:
- AI Platform Strategy & Architecture:
Assess, define and evolve enterprise AI platform architectures, covering LLM and agentic frameworks, AI gateways, model lifecycle management, data platforms, MLOps/LLMOps foundations and integration patterns. Support clients in evaluating build, buy and hybrid approaches aligned to business needs, risk appetite and operational requirements. - AI Platform Engineering & LLMOps:
Design and implement scalable AI platform capabilities including model deployment pipelines, prompt and model management, evaluation frameworks, AI observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production. - Reliability Engineering & SRE:
Establish SRE practices including SLIs, SLOs, error budgets, capacity planning, resilience engineering and reliability governance. Help clients shift from reactive operations to data‑driven reliability management while balancing reliability, innovation and delivery velocity. - Observability & Operational Intelligence:
Define observability strategies across applications, platforms and AI workloads using metrics, logs, traces and telemetry. Establish operational insight models that support proactive decision‑making and enable advanced capabilities including anomaly detection, event intelligence, noise reduction and predictive operational analytics. - AI Operations & Service Reliability:
Apply reliability engineering principles to AI‑enabled services, monitoring AI‑specific failure modes such as data quality degradation, hallucination patterns, token consumption, agent reliability and model performance drift. Implement controls, feedback loops and automated guardrails to ensure AI services remain secure, trusted and cost‑effective. - Responsible AI & Platform Governance:
Design and embed AI governance frameworks, model risk controls, compliance measures and responsible AI practices that address security, regulatory and ethical requirements while supporting innovation and adoption at scale. - Automation & Operational Efficiency:
Identify opportunities to reduce operational complexity and toil through engineering‑led automation, intelligent workflows and AI‑enhanced operational practices. Help clients improve scalability, consistency and operational performance while reducing manual effort. - Client Advisory & Transformation Leadership:
Act as a trusted advisor to CIO, CTO, CDO and Engineering leadership stakeholders, shaping platform strategies, operating models, vendor selections and transformation roadmaps. Lead consulting teams and work streams from assessment and strategy through implementation and scale‑up.
- Proven experience designing, delivering and operating cloud‑native, platform engineering, AI platform or reliability engineering solutions…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: