×
Register Here to Apply for Jobs or Post Jobs. X

Lead AI Infrastructure Operations Engineer

Remote / Online - Candidates ideally in
Iowa, Calcasieu Parish, Louisiana, 70647, USA
Listing for: The Mutual Group
Remote/Work from Home position
Listed on 2026-09-06
Job specializations:
  • IT/Tech
    SRE/Site Reliability, IT Infrastructure, Cloud Computing: Infrastructure & Operations, AI Engineer (Applied/Software)
Salary/Wage Range or Industry Benchmark: 130000 - 150000 USD Yearly USD 130000.00 150000.00 YEAR
Job Description & How to Apply Below
Location: Iowa

Department:
Information Technology

Job Description:

The Lead AI Infrastructure Operations Engineer is responsible for enabling, operating, and continuously improving the infrastructure and operational capabilities required by production AI solutions. As TMG expands its AI capabilities, AI CoE will develop large language model, retrieval-augmented generation, agent-based, machine-learning, and other AI-enabled use cases. This role will work within the Infrastructure team to help those teams define and coordinate their infrastructure, environment, connectivity, observability, security, capacity, and production-readiness needs.

The engineer in collaboration with other Infrastructure resources will focus on the AI application, platform, and operational layers. TMG’s managed services provider operates the underlying AWS cloud infrastructure. This role will translate AI solution requirements, coordinate infrastructure services and changes, monitor delivery, diagnose issues, and validate that environments meet reliability, security, performance, and operational expectations. This is a hands-on role that works closely with AI engineering, application development, data, architecture, cybersecurity, infrastructure operations, and external managed-services partners.

Work

Arrangement

Employees who live within 30 miles of the TMG home office are expected to follow a hybrid or in-office schedule. The initial training period may require additional in-office days.

Accountabilities

Enable AI Infrastructure and Environments Partner with AI engineering, application, data, security, architecture, and platform teams to understand the infrastructure and operational needs of new AI use cases. Translate AI solution designs into requirements for environments, compute, storage, networking, connectivity, identity, security, observability, capacity and supporting cloud services. In collaboration with other Infrastructure resources and managed service provider, coordinate infrastructure provisioning, configuration, access and changes with TMG’s AWS managed-services provider and other technology partners.

Support operational readiness for AI solutions transitioning into production. Identify infrastructure dependencies, constraints, risks, costs and lead times early in the delivery lifecycle. Establish reusable infrastructure patterns, operational standards, dashboards, runbooks and production-readiness requirements across AI use cases. Validate that environments and supporting services are appropriately configured, monitored, secured, scalable and ready for production use. Implement new cloud functionality requirements in collaboration with architecture and AI engineering within approved architecture and guardrails.

Follow and enforce established security, compliance and operational controls across AI platforms and infrastructure.

Operate and Improve Production AI Solutions Work with managed services to implement and maintain observability, logging, tracing, monitoring, dashboards and alerting for production AI applications. Define and track operational metrics covering AI quality, reliability, latency, cost, utilization, adoption and business outcomes. Work with AI engineering team to diagnose production issues across AI applications, models, prompts, retrieval systems, data, integrations and supporting infrastructure.

Support root‑cause analysis and coordinate resolution with AI engineers, application teams, platform teams and managed‑services providers. Support AI‑related incident management, problem management, operational reviews, release validation and production-readiness activities. Create and maintain service‑health dashboards, runbooks, troubleshooting guidance, support procedures and escalation paths. Define and track service health indicators, SLIs, SLOs and operational KPIs for production AI solutions. Use telemetry, evaluations and production data to verify fixes, releases, configuration changes and system improvements.

Support AI Risk and Operational Governance Partner with AI and IT governance team to operationalize applicable controls and monitoring requirements. Operationalize evaluation processes and…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary