×
Register Here to Apply for Jobs or Post Jobs. X

Lead AI Infrastructure Operations Engineer

Remote / Online - Candidates ideally in
Iowa, USA
Listing for: The Mutual Group
Remote/Work from Home position
Listed on 2026-09-04
Job specializations:
  • IT/Tech
    SRE/Site Reliability, AI Engineer (Applied/Software), IT Infrastructure, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below

Lead AI Infrastructure Operations Engineer

The Lead AI Infrastructure Operations Engineer is responsible for enabling, operating, and continuously improving the infrastructure and operational capabilities required by production AI solutions.

As TMG expands its AI capabilities, AI CoE will develop large language model, retrieval-augmented generation, agent-based, machine-learning, and other AI-enabled use cases. This role will work within the Infrastructure team to help those teams define and coordinate their infrastructure, environment, connectivity, observability, security, capacity, and production-readiness needs.

The engineer in collaboration with other Infrastructure resources will focus on the AI application, platform, and operational layers. TMG's managed services provider operates the underlying AWS cloud infrastructure. This role will translate AI solution requirements, coordinate infrastructure services and changes, monitor delivery, diagnose issues, and validate that environments meet reliability, security, performance, and operational expectations.

This is a hands-on role that works closely with AI engineering, application development, data, architecture, cybersecurity, infrastructure operations, and external managed-services partners.

Work Arrangement:

  • Employees who live within 30 miles of the TMG home office are expected to follow a hybrid or in-office schedule. The initial training period may require additional in-office days.
Accountabilities
  • Enable AI Infrastructure and Environments
    • Partner with AI engineering, application, data, security, architecture, and platform teams to understand the infrastructure and operational needs of new AI use cases.
    • Translate AI solution designs into requirements for environments, compute, storage, networking, connectivity, identity, security, observability, capacity, and supporting cloud services.
    • In collaboration with other Infrastructure resources and managed service provider, coordinate infrastructure provisioning, configuration, access, and changes with TMG's AWS managed-services provider and other technology partners.
    • Support operational readiness for AI solutions transitioning into production.
    • Identify infrastructure dependencies, constraints, risks, costs, and lead times early in the delivery lifecycle.
    • Establish reusable infrastructure patterns, operational standards, dashboards, runbooks, and production-readiness requirements across AI use cases.
    • Validate that environments and supporting services are appropriately configured, monitored, secured, scalable, and ready for production use.
    • Implement new cloud functionality requirements in collaboration with architecture and AI engineering within approved architecture and guardrails.
    • Follow and enforce established security, compliance, and operational controls across AI platforms and infrastructure.
  • Operate and Improve Production AI Solutions
    • Work with managed services to implement and maintain observability, logging, tracing, monitoring, dashboards, and alerting for production AI applications.
    • Define and track operational metrics covering AI quality, reliability, latency, cost, utilization, adoption, and business outcomes.
    • Work with AI engineering team to diagnose production issues across AI applications, models, prompts, retrieval systems, data, integrations, and supporting infrastructure.
    • Support root-cause analysis and coordinate resolution with AI engineers, application teams, platform teams, and managed-services providers.
    • Support AI-related incident management, problem management, operational reviews, release validation, and production-readiness activities.
    • Create and maintain service-health dashboards, runbooks, troubleshooting guidance, support procedures, and escalation paths.
    • Define and track service health indicators, SLIs, SLOs, and operational KPIs for production AI solutions.
    • Use telemetry, evaluations, and production data to validate fixes, releases, configuration changes, and system improvements.
  • Support AI Risk and Operational Governance
    • Partner with AI and IT governance team to operationalize applicable controls and monitoring requirements.
    • Operationalize evaluation processes and support AI Governance team to monitor AI performance, regressions, drift, grounding, retrieval quality, and overall effectiveness.
    • Support the collection and retention of operational evidence, including model and prompt versions, evaluation results, incidents, exceptions, and corrective actions.
    • Identify material changes in AI behavior and help ensure they are evaluated, documented, and appropriately addressed.
    • Analyze trends and proactively identify degradation, drift, capacity constraints, reliability risks, and quality issues before they become production incidents
Key Outcomes
  • AI teams receive timely and consistent infrastructure and operational support.
  • AI use cases move efficiently from experimentation to reliable production operation.
  • Reusable infrastructure and operational patterns are applied across AI initiatives.
  • End-to-end visibility…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary