More jobs:
AI Infrastructure Operations Engineer
Job in
Minneapolis, Hennepin County, Minnesota, 55405, USA
Listed on 2026-09-04
Listing for:
Accenture
Full Time
position Listed on 2026-09-04
Job specializations:
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure, SRE/Site Reliability
Job Description & How to Apply Below
We embrace the power of change to create value and shared success for our clients, people, shareholders, partners, and communities. Visit us at
The Global AI Infrastructure team enables resilient, high-performance compute environments for strategic clients across cloud, on-premises, and hybrid deployments. We design, build, and operate large-scale GPU and accelerated-computing infrastructure that supports demanding AI training and inference, simulation, and high-performance compute workloads. Our work spans strategy, architecture, modernization, operations, governance, and continuous improvement across the infrastructure stack. We build reusable operational tools, automation workflows, and platform capabilities that make repeatable infrastructure tasks safer, faster, and more scalable.
We collaborate across the technology ecosystem to harness new capabilities, drive business transformation, and deliver dependable services at scale.
Key Responsibilities:
+ Design and implement accelerated-computing infrastructure solutions aligned to system architecture, deployment roadmaps, performance, scalability, resiliency, and governance requirements.
+ Deploy, configure, and operate GPU-based clusters across bare-metal and containerized environments, using workload schedulers and Kubernetes orchestration to support AI training, inference, and high-performance compute workloads.
+ Integrate infrastructure platforms with enterprise systems, data platforms, security frameworks, service-management processes, and governance controls.
+ Design, build, and maintain reusable tools, scripts, self-service capabilities, and automation workflows for infrastructure operations, including provisioning, configuration management, validation, capacity planning, monitoring, incident management, reporting, and recurring remediation.
+ Establish repeatable operational processes for cluster provisioning, configuration management, patching, capacity planning, monitoring, incident response, and lifecycle management.
+ Perform and automate GPU, compute, storage, and network benchmarking and validation; diagnose performance issues across multi-node AI training, inference, and distributed compute workloads.
+ Develop and maintain architecture diagrams, configuration baselines, operational runbooks, and support documentation.
+ Provide technical guidance, troubleshooting, and optimization for GPU clusters supporting AI training, inference, high-performance computing, and multi-node simulation workloads, with emphasis on availability, resiliency, scalability, energy efficiency, and cost management.
Travel may be required for this role. The amount of travel will vary from 25% to 60% depending on business need and client requirements.
Required Skills and Qualifications:
+ Minimum of 5+ years of experience designing, deploying, and managing accelerated-computing infrastructure across on-premises, cloud, and hybrid environments for hyperscaler, neocloud, large enterprise, telecommunications, financial services, manufacturing, and/or retail clients.
+ Minimum of 5+ years of hands-on experience with accelerated-computing platforms, including GPUs, DPUs, and CPUs, high-bandwidth network fabrics, and AI based storage architectures such as parallel file systems, NVMe-oF, etc.
+ Minimum of 5+ years of experience with cluster management, workload scheduling, orchestration, observability, and infrastructure automation, including building operational tools and automation workflows with platforms such as Kubernetes, Slurm, Run:ai
+ Minimum 6 months hands-on experience with Claude Code, AI automation tools, Terraform, Ansible, Python, and Bash scripting.
+ Bachelor's degree or equivalent (minimum 12 years) work experience. (If Associate's Degree, must have minimum 6 years work experience)
Preferred Skills and
Qualifications:
+ Experience building AI infrastructure automation and operations tools, Agentic Ops practices that enable secure, automated, governed, and reproducible platform operations.
+ Experience developing reusable infrastructure code leveraging Python, platform services, and automation workflows using REST APIs, OpenAPI, JSON/YAML schemas, webhooks, and event-driven integrations.
+ Experience operating large-scale GPU clusters, including capacity management, reliability engineering, change management, and performance validation for AI training, inference, HPC, and enterprise compute workloads.
+ Experience using NVIDIA platform tools and libraries including Base…
Position Requirements
5+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×