×
Register Here to Apply for Jobs or Post Jobs. X

Senior DevOps Engineer

Job in Toronto, Ontario, C6A, Canada
Listing for: MarkiTech.AI
Full Time position
Listed on 2026-08-23
Job specializations:
  • IT/Tech
    AWS, SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Salary/Wage Range or Industry Benchmark: 140000 - 180000 CAD Yearly CAD 140000.00 180000.00 YEAR
Job Description & How to Apply Below

Senior Dev Ops Engineer (alternate: Senior Cloud / Platform Engineer)

We are hiring a senior Dev Ops engineer to own and evolve our cloud platform on AWS, grounded in infrastructure as code, secure multi-account patterns, and reliable delivery. You will shape the Dev Ops roadmap (standards, tooling, automation, and operational excellence), support application releases, and provide production support for critical workloads.

Amazon EKS is central to how we run workloads—we need someone with deep, production-grade EKS expertise who has built and owned Kubernetes on AWS end-to-end, not only deployed apps to a shared cluster someone else runs.

You will also lead how we adopt AI for infrastructure and platform work—not as a buzzword, but as a practical force multiplier: safe use of AI-assisted authoring and review for IaC and automation, clearer runbooks and incident workflows, and evaluation of tools and patterns that improve speed without weakening security, compliance, or change control. This role suits someone who combines deep AWS practice with leadership: you can define "how we build and run" while still being hands‑on in pipelines, clusters, and incidents.

What

You Will Do
  • Roadmap & standards:
    Define and socialize Dev Ops priorities (security, reliability, cost, velocity). Align teams on AWS Well-Architected practices, tagging, guardrails, and repeatable patterns for networking, identity, secrets, and data.
  • AI adoption for infra & platform:
    Drive a pragmatic AI strategy for the team—e.g. standards for AI-assisted IaC and pipeline changes (review gates, testing, drift detection), documentation and runbook quality, incident summarization and triage workflows where appropriate, and guardrails so AI tooling fits regulated or high-stakes environments. Stay current on vendor and open-source options; pilot, measure, and roll out what actually reduces toil.
  • Infrastructure as code:
    Design, review, and implement changes using Terraform and Terragrunt, with clear module boundaries, environment‑specific config, and safe promotion across dev, non‑prod, and production.
  • EKS (critical):
    Build, operate, and own the Kubernetes platform on AWS —cluster lifecycle (creation, upgrades, patching), node groups / capacity, networking (CNI, service mesh or ingress as used), security (RBAC, admission controls, pod security, secrets and IRSA), add‑ons, and cost/ reliability tuning. Partner with app teams on standards for workloads, name spaces, and safe rollouts; be the escalation point for cluster-level incidents.
  • Broader AWS platform:
    Operate and improve adjacent services—e.g. RDS/Aurora, DynamoDB, object storage and CDN, KMS, Secrets
  • Manager, SNS (alerting), Lambda, Event Bridge, and CI/CD (Code Pipeline / Code Build, connections to source control)—plus IAM, VPC, and multi‑tenant or multi‑namespace patterns where applicable.
  • Release engineering:
    Partner with development teams on release processes, deployment strategies, change management, rollbacks, and post-release verification in regulated or high-stakes environments (e.g. healthcare‑adjacent workloads).
  • Production support:
    Participate in on‑call or escalation rotation as defined by the team; troubleshoot incidents, drive root‑cause analysis, and implement preventive fixes (runbooks, dashboards, alarms, automation).
  • Observability & operations:
    Improve monitoring, logging, tracing, and alerting; tune thresholds; reduce noise; document operational procedures.
  • Collaboration:

    Work with security, architecture, and engineering leads to implement least‑privilege access, encryption, backup/DR posture, and audit‑friendly operations—including how AI‑assisted workflows meet security and audit expectations.
Required

What we are looking for:

  • 6+ years in software / systems / Dev Ops / SRE roles, including 4+ years focused on AWS in production.
  • Strong command of infrastructure as code (Terraform) and modular, environment-driven layouts (experience with Terragrunt or similar composition patterns is a plus).
  • Deep, mandatory expertise in Amazon EKS:
    You have prior experience building and owning Kubernetes on AWS—not only deploying applications to a shared cluster. We expect fluency across the…
Position Requirements
10+ Years work experience
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary