Main Activities
AI Platform Reliability and Operations
• Administer and operate the AI platform to ensure availability, performance, and resilience across environments, integrations, and supporting infrastructure.
• Monitor platform health using dashboards, logs, metrics, and alerts, and coordinate incident and service restoration activities.
• Lead operational triage, escalation coordination, and post-incident reviews to strengthen services stability and resilience.
• Track and report on service reliability indicators, incident trends, and operational performance.
AI Platform Enablement & Production Readiness
• Enable approved AI use cases into production by ensuring:
o Environment readiness,
o Dependency validation,
o Completion of operational readiness checklists,
o Structured service transition activities
• Support platform lifecycle management through:
o Release coordination,
o Change readiness validation,
o Maintenance and capacity planning.
• Ensure AI platform changes meet defined operational and control readiness criteria prior to release
Observability, Automation & AI Ops
• Implement and maintain observability capabilities, including telemetry, logging, metrics, and traces required for enterprise AI operations.
• Analyze operational data to identify anomalies, recurring issues, root-cause patterns.
• Implement AI Ops use cases such as:
o Alert correlation,
o Anomaly detection,
o Root-cause support,
o Forecasting and predictive insights,
o Automation of repetitive operational tasks.
• Continuously improve operational efficiency through targeted automation and process optimization.
Governance, Risk, & Control Execution
• Execute governance controls for AI solutions, including:
o Usage and access controls,
o Data privacy considerations,
o Auditability and traceability,
o Human oversight requirements
• Ensure operational practices align with enterprise security policies, risk controls, and compliance requirements.
• Maintain documentation and evidence required for audit, governance reviews, production readiness checkpoints, and control validation.
• Identify control gaps and escalate risks appropriately to relevant governance and risk stakeholders.
AI Asset Visibility & Operational Integrity
• Maintain operational visibility of AI platform assets required for monitoring, support, and cost alignment.
• Validate asset ownership, relationships, and lifecycle status in collaboration with application and platform owners.
• Support ongoing audits to ensure AI assets and associated cost attribution remain accurate and current.
Knowledge/Skill Requirements
• University degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience.• 5-7 years of experience in platform operations, site reliability engineering, Dev Ops, cloud operations, or enterprise IT operations.
• Strong experience supporting production platforms and services, including monitoring, incident response, problem management, service restoration, and operational reporting.
Technical Expertise:
• Experience with cloud platforms, observability, automation, configuration management, and integration patterns, including Azure Automation runbooks (Power Shell/Python), Azure AI, Copilot integrations, AKS, virtual networks (hub-and-spoke), and App Service.
• Expertise with observability tools such as Azure Monitor, Application Insights, and Grafana.
• Experience with CI/CD and automation tools such as Azure Dev Ops, Git Hub Actions, and Logic Apps.
• Knowledge of configuration management and infrastructure-as-code tools such as Bicep, Terraform, Azure Policy, Key Vault, and relevant open-source technologies.
• Knowledge of integration and event-driven technologies such as API Management, open-source API tools, Service Bus, Event Grid, and Apache Kafka.
• Working knowledge of platform-supporting data and search services such as Elastic, Azure AI Search, and Cosmos DB.
• Knowledge of enterprise network, edge security, and related internal platforms such as DNA, Fortinet, and Akamai is an asset.
Additional Capabilities:
• Working knowledge of AI/ML operational concepts, including model lifecycle support, telemetry, governance…
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search: