×
Register Here to Apply for Jobs or Post Jobs. X

AI Systems Engineer - AI Platforms - Manager

Job in Indianapolis, Marion County, Indiana, 46202, USA
Listing for: EY
Full Time position
Listed on 2026-08-30
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Job Description & How to Apply Below
Location:

Anywhere in Country

At EY, we're all in to shape your future with confidence.

We'll help you succeed in a globally connected powerhouse of diverse teams and take your career wherever you want it to go.  Join EY and help to build a better working world.

** The opportunity*
* We are seeking AI Systems Engineers to build and operate the foundational substrate that powers EY's AI-native platform. This role owns the infrastructure and cloud-native platform layers of the Hybrid AI Multi-Environment Runtime (HAI), from bare-metal and GPU infrastructure through Kubernetes, cluster fabric, and multi-tenant scaling. You will be responsible for building and managing EY Fabric environments across cloud, on-prem, edge, and air-gapped targets.

This is the substrate on which EY Agentic AI capabilities run. This role is ideal for a full-stack infrastructure leader who is equally comfortable with bare-metal and GPU systems and production Kubernetes at scale, who treats reliability and portability as non-negotiable in regulated client contexts, and who understands that the substrate is a product in its own right, measured by the velocity, safety, and portability it unlocks for every team building above it.

** Your key responsibilities*
* +  
** Own the cluster & cloud-native platform:
** compute, Kubernetes and scheduling, cluster fabric/networking, multi-tenancy, and distributed compute, as the substrate for Agentic AI workflows and tooling.

+  
** Own the infrastructure foundation:
** Ubuntu/OS, BMC/bare-metal, DPU architecture, and NVAIE (GPU/Network/DCGM), ensuring the physical and virtual bedrock is provisioned, patched, and production-ready.

+ Stand up and manage EY Agentic AI environments across cloud (EKS/AKS/GKE), on-prem AI Factory (RKE2/NVAIE), edge (K3s), and air-gapped deployment modes, maintaining one consistent stack contract across all targets.

+ Deliver foundational platform capabilities such as Infrastructure Management, Kubernetes & Scheduling, and Cluster Fabric Management, so downstream runtime, data, and execution services can run safely and consistently.

+ Own cluster lifecycle, autoscaling, GPU pooling/virtualization, and multi-tenancy boundaries (vCluster/Crossplane/Karpenter), providing isolated, elastic capacity per tenant and engagement.

+  
** Own secure execution and inference:
** Ray Serve, vLLM/NIM/Triton, and NVIDIA Dynamo, with sandboxed execution (gVisor/Firecracker for hosted, NVIDIA Open Shell/vNode for on-prem) for isolated, safe model execution.

+  
** Own cognitive and routing:
** Envoy AI Gateway, semantic routing (vLLM-SR), model/prompt selection, and streaming response handling - directing each request to the right model under the right constraints.

+ Collaborate with Dev Ops Engineers on deployment and delivery of the platform itself: CI/CD/CV (ArgoCD), infrastructure-as-code / Git Ops (Helm/Open Tofu), so environments are reproducible and drift-free.

+ Own backup, disaster recovery, and cross-environment replication for high availability (Velero, CloudNative PG, Cilium Cluster Mesh), along with patching and platform supply-chain hygiene.

+ Ensure the substrate is modular and swappable, so components can be replaced without rewriting consumers, minimizing vendor lock-in while preserving the stack contract.

** Skills and attributes for success*
* + Deep expertise in cloud-native platform engineering (Kubernetes, networking, multi-tenancy at scale) and low-level infrastructure and systems engineering (bare-metal, GPU, DPU).

+ Advanced understanding of how compute, networking, storage, and scheduling interact to form a reliable, portable substrate.

+ Comfortable operating across cloud, on-prem, edge, and air-gapped environments simultaneously, with a portability-first mindset.

+ Strong command of AI gateways, semantic routing, and model/prompt selection under latency and cost constraints.

+ A passion for ensuring reliability, repeatability, and reduction of operational toil through automation and infrastructure-as-code.

+ Ability to define and honor clean ownership boundaries with adjacent trust, data, and runtime teams.

+ Strong communicator able to explain infrastructure tradeoffs to engineers, architects, and leadership.

+ Natural product-ownership orientation toward foundational platforms, measured by the downstream velocity and reliability they enable.

** To qualify you must have*
* + Bachelor's or Master's degree in Computer Science or related technical field.

+ 8+ years building or operating enterprise infrastructure, cloud platforms, or large-scale Kubernetes environments, including hands-on systems depth.

+ Hands-on expertise with Kubernetes distributions (RKE2, EKS/AKS/GKE, K3s) and full cluster lifecycle management.

+ Deep experience with bare-metal, cloud, hybrid, on-prem, and ideally air-gapped deployment models.

+ Strong grounding in cluster networking (Cilium/service mesh/CNI), storage, and multi-tenancy isolation.

+

Experience with GPU infrastructure and scheduling (NVAIE/DCGM, GPU operators,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary