×
Register Here to Apply for Jobs or Post Jobs. X

Director, AI Operations – Optimization

Job in Springfield, Sangamon County, Illinois, 62777, USA
Listing for: Jobtailor
Full Time position
Listed on 2026-07-20
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 150000 - 210000 USD Yearly USD 150000.00 210000.00 YEAR
Job Description & How to Apply Below

Responsibilities

  • Lead enterprise AI operational activities, including runtime monitoring, operational support, incident management, production reliability, and operational continuity for AI-powered applications and intelligent automation solutions.
  • Establish, implement, and continuously improve AI operational practices, including AIOps and LLMOps processes, runtime observability, operational telemetry, drift detection, release coordination, support workflows, and operational readiness activities.
  • Drive runtime stability and service reliability initiatives through production monitoring, escalation management, root cause analysis, operational playbooks, and service continuity practices.
  • Support enforcement of runtime governance standards, operational safeguards, human oversight controls, and secure operationalization practices for enterprise AI solutions.
  • Ensure operational excellence across AI environments through proactive monitoring, issue prevention, and continuous service improvement efforts.
  • Lead initiatives focused on runtime efficiency, operational scalability, inference utilization, supportability, performance optimization, and sustainable AI operations.
  • Promote standardized operational processes, scalable support models, automation opportunities, and continuous improvement initiatives across AI operations functions.
  • Drive operational maturity by identifying opportunities to enhance performance, reduce operational risk, and improve support effectiveness.
  • Partner closely with AI Engineering & Delivery, AI Governance, AI Quality Engineering, Automation, Architecture, Platform Engineering, Security, Infrastructure, and business stakeholders to ensure operational readiness and runtime reliability.
  • Coordinate operational execution activities across AI operations teams, including operational planning, vendor and contractor management, issue prioritization, escalation management, knowledge transfer, and delivery continuity.
  • Support operational assessments, production readiness reviews, implementation planning, runtime support strategies, and modernization initiatives for prioritized AI capabilities.
  • Collaborate with technical and business leaders to align operational practices with enterprise AI objectives and service expectations.
  • Lead, mentor, and develop operations managers, engineers, analysts, and contractor resources while fostering a high-performing, collaborative, and continuously learning culture.
  • Provide clear communication regarding operational performance, runtime risks, service reliability concerns, optimization opportunities, engineering tradeoffs, and strategic recommendations.
  • Establish accountability for operational outcomes while promoting operational discipline, innovation, and continuous improvement.
  • Research and evaluate emerging AI operational technologies, observability platforms, automation capabilities, optimization techniques, and runtime management practices to drive innovation and operational effectiveness.
Requirements
  • Bachelor’s degree in Computer Science, Information Systems, Engineering, Technology Management, or a related field preferred.
  • 8+ years of experience in AI operations, software engineering, platform operations, engineering delivery, Dev Ops, Site Reliability Engineering (SRE), infrastructure operations, or related enterprise technology functions required.
  • 3+ years of experience leading operational teams, engineering support organizations, platform operations, or large-scale technology initiatives required.
  • Hands‑on experience supporting, operationalizing, monitoring, or optimizing production AI solutions utilizing large language models (LLMs), APIs, agentic workflows, orchestration frameworks, and modern AI engineering practices required.
  • Strong experience implementing and scaling operational support models, observability practices, incident management processes, Dev Ops methodologies, runtime operations, or enterprise operational frameworks required.
  • Experience with observability platforms, monitoring tools, incident management processes, runtime operations, CI/CD pipelines, and production support practices required.
  • Experience leading distributed teams, managing…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary