AI Platform Operations Manager
Listed on 2026-08-30
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Infrastructure, Azure
The Company
STACK INFRASTRUCTURE (STACK) provides digital infrastructure to scale the world’s most innovative companies. We are an award-winning industry leader in building, owning, and operating highly efficient, cost-effective wholesale, colocation, and cloud data centers. Each of our national facilities meets or exceeds the highest industry standards in all operational categories of availability, security, connectivity, and physical resilience.
The CompanySTACK INFRASTRUCTURE (STACK) provides digital infrastructure to scale the world’s most innovative companies. We are an award-winning industry leader in building, owning, and operating highly efficient, cost-effective wholesale, colocation, and cloud data centers. Each of our national facilities meets or exceeds the highest industry standards in all operational categories of availability, security, connectivity, and physical resilience.
STACK offers the scale and geographic reach that rapidly growing hyperscale and enterprise companies need. The world runs on data. Data runs on STACK.
The PositionThe Dev Ops Engineer, AI Platform is responsible for automating, deploying, and operating the infrastructure and delivery pipelines that support STACK’s enterprise AI platform on Azure. This is a hands‑on engineering role focused on build and run — not oversight. Reporting to Head of AI, Enterprise AI & Data Strategy org this individual owns the infrastructure-as-code, CI/CD, containerization, observability, and release automation that allow AI engineers and enterprise application teams to ship agentic AI solutions, RAG pipelines, and integration services reliably and repeatably.
The role sits at the intersection of cloud infrastructure, platform engineering, and MLOps — turning platform architecture into automated, governed, observable, and cost‑efficient environments that teams across the organization build on.
- Build, maintain, and version infrastructure-as-code modules for Azure environments using Terraform, Bicep, or ARM, including compute, networking, storage, identity, and AI platform resources.
- Automate provisioning of AI platform components — Azure AI Foundry, Azure OpenAI Service, Azure AI Search, Cosmos DB, ADLS Gen2, and Databricks — as reusable, parameterized deployment patterns.
- Maintain environment parity across development, test, and production, including configuration management, drift detection, and remediation.
- Implement and enforce tagging, naming, and resource organization standards that support governance, chargeback, and lifecycle management.
- Automate routine platform operations — patching, certificate rotation, key and secret rotation, backup validation, and disaster recovery testing.
- Design, build, and operate CI/CD pipelines in Azure Dev Ops or Git Hub Actions for application code, infrastructure code, container images, and AI/agent deployments.
- Implement automated build, test, security scanning, artifact management, and promotion gates across environments.
- Establish branching strategies, code review standards, and release management practices in partnership with AI engineering and enterprise application teams.
- Build deployment automation for agentic AI services, MCP (Model Context Protocol) servers, and integration workloads running on Azure Container Apps and Azure Kubernetes Service (AKS).
- Support model and prompt release workflows — versioning, staged rollout, evaluation gates, and rollback procedures for LLM-based applications.
- Operate and tune AKS and Azure Container Apps, including cluster upgrades, node pool sizing, autoscaling, ingress, networking, and workload isolation.
- Build and maintain container images, base image standards, and registry governance in Azure Container Registry.
- Manage compute scheduling and scaling for AI workloads, including GPU-backed and inference-heavy workloads where required.
- Implement resiliency patterns — health probes, retries, throttling, quota management, and failover — for AI endpoints and integration services.
- In…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).