Senior IT Application Specialist, DevOps & AI Operations - SaaS Client
Job in
Regina, Saskatchewan, S4M, Canada
Listing for:
S.i. Systems
Full Time
position
Listed on 2026-10-10
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Senior IT Application Specialist, Dev Ops & AI Operations
- SaaS Client
Type: FTE/Permanent
Location: Oakville, ON
- Hybrid if local
Overview
Our client is seeking a Senior IT Application Specialist with a strong Dev Ops background to manage enterprise applications and cloud platforms s hands-on role combines Kubernetes platform operations, infrastructure automation and observability with daily AI use to improve monitoring, log investigation, troubleshooting and operational efficiency.
Responsibilities
Own the lifecycle of enterprise applications and supporting platforms, including design, deployment, ongoing support, performance tuning and upgrades.Manage containerized workloads using Docker, Kubernetes and Helm, with Terraform for infrastructure automation and Git Ops/ArgoCD for application delivery.Maintain monitoring and logging using Grafana, Elasticsearch and Victoria Metrics; automate health checks and improve incident detection and investigation.Use AI tools in daily operations to assist with coding, log analysis, troubleshooting, documentation and workflow automation.Evaluate and implement AI tools and integrations with appropriate security, testing and governance.Lead complex incident resolution, collaborate with business and technical teams, mentor colleagues, and maintain operational documentation and runbooks.Must Haves
5–8 years of relevant experience managing production cloud platforms and/or enterprise applications, with substantial hands-on Dev Ops responsibility.Experience with Docker, Kubernetes and Helm
, including deploying, configuring and troubleshooting production workloads.Hands-on experience with Terraform, Git Ops and ArgoCD for infrastructure provisioning and application deployment.Experience with Grafana, Elasticsearch and Victoria Metrics for monitoring, logging and platform health.Production cloud experience with Google Cloud or AWS
. Strong AWS experience will be considered transferable.Full willingness to adopt AI in day-to-day operations
, including learning and applying AI tools to improve investigation, automation and support.Expert-level Linux administration and strong scripting/programming skills in Python, Bash and Power Shell
.Strong knowledge of SSO methodologies, including SAML and LDAPS, and Zero Trust security principles
.Ability to troubleshoot complex production issues, communicate clearly with technical and business stakeholders, and take ownership through resolution.Availability to participate in a 24×7 rotating on-call schedule
.Relevant post-secondary education or an equivalent combination of education and demonstrable technical experience.Nice to Haves
Direct experience with Google Cloud Platform and Google Kubernetes Engine (GKE).Existing daily use of Claude Code, Codex, Open Code, Gemini CLI or comparable AI tools for coding and operational work.Practical experience using AI for monitoring, log investigation, anomaly detection, root-cause analysis or incident response
.Experience building workflow automations and AI integrations using LLM APIs, n8n, GCP Cloud Functions or AWS Lambda
, including governed agent workflows.
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here: