Applied AI Site Reliability Engineer III - PxE Talent
Listed on 2026-08-03
-
Software Development
AI Engineer (Applied/Software)
Applied AI Site Reliability Engineer III
Role Overview:
As an Applied AI Site Reliability Engineer III , you will actively engage in your engineering craft, taking a hands‑on approach to the reliability, performance, and operational integrity of high‑visibility products and platforms and the environments they run in. Your expertise will be pivotal in keeping production safe, performant, and cost‑effective, while driving tangible value for Deloitte's engineering investments. You will leverage your extensive engineering craftsmanship across cloud platform engineering, observability, and performance and reliability engineering‑together with applied AI fluency that lets you reliably operate AI and agentic workloads alongside the rest of the portfolio‑consistently demonstrating your strong track record in operating high‑quality, resilient systems ideal candidate will be a dependable team player, collaborating with cross‑functional teams to uphold production standards, safeguard environments, and admit systems into production with confidence.
team:
US Deloitte Technology Product Engineering has modernised software and product delivery, creating a scalable, cost‑effective model that focuses on value/outcomes that leverages a progressive and responsive talent structure. As Deloitte's primary internal development team, Product Engineering delivers innovative digital solutions to businesses, service lines, and internal operations with proven bottom‑line results and outcomes. It helps power Deloitte's success. It is the engine that drives Deloitte, serving many of the world's largest, most respected companies.
We develop and deploy cutting‑edge internal and go‑to‑market solutions that help Deloitte operate effectively and lead in the market. Our reputation is built on a tradition of delivering with excellence.
- Excellent interpersonal and organizational skills, with the ability to handle diverse situations, complex projects, and changing priorities, behaving with passion, empathy, and care.
- A bachelor's degree in computer science, software engineering, data science, machine learning, or related discipline. Experience is the most relevant factor.
- 5+ years of software engineering and site reliability engineering experience operating large‑scale, distributed, cloud‑native systems in production, with experience in most of the following:
Python, Go, Bash, Java, C#/.NET, SQL/No
SQL, Kubernetes, Terraform, ArgoCD, as well as CI/CD and observability stacks. - 3+ years of experience in site reliability or production engineering for large‑scale systems‑defining and owning SLIs, SLOs, and SLAs; error budgets; incident command and on‑call; building and operating production observability (metrics, tracing, logging‑e.g., Open Telemetry, Prometheus, Grafana, Datadog, Dynatrace, Amazon Cloud Watch, Azure Monitor, Google Cloud Operations, Solar Winds, Splunk); environment integrity and drift prevention across pre‑production and production; and segregation‑of‑duties controls (least‑privilege/RBAC, deploy approvals, secrets management) in partnership with security and risk.
- 3+ years of experience with cloud‑native engineering and cloud platform ownership on any of the cloud hyperscalers such as Azure, AWS, or GCP‑including their AI/ML services such as Azure OpenAI, AWS Bedrock, or Vertex AI‑plus container orchestration (Kubernetes, Docker), infrastructure‑as‑code, networking, and multi‑environment management.
- Prior experience operating AI/ML and agentic workloads in production‑their reliability failure modes (drift, train/serve skew, output variance), MLOps/LLMOps, and the AI control plane (model/LLM gateway, guardrails) from the operability and performance side.
- Prior experience with load and performance testing under simulated production traffic (e.g., Load Runner, k6, or JMeter), chaos engineering (e.g., Azure Chaos Studio, AWS Fault Injector), capacity planning, autoscaling, and cloud/AI cost engineering (Fin Ops tooling/dashboards, including GPU/inference and token cost attribution).
- Prior software engineering experience with the understanding of Business…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).