×
Register Here to Apply for Jobs or Post Jobs. X

Lead Platform Engineer; DevOps & MLOps

Job in Toronto, Ontario, C6A, Canada
Listing for: SmartRecruiters, Inc.
Full Time position
Listed on 2026-10-07
Job specializations:
  • IT/Tech
    SRE/Site Reliability, IT Infrastructure, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 120000 - 180000 CAD Yearly CAD 120000.00 180000.00 YEAR
Job Description & How to Apply Below
Position: Lead Platform Engineer (DevOps & MLOps)

Shore is an IT and strategy consulting firm focusing on innovation in the public sector. We deliver services and tools that advance public sector organizations and the services they provide.

Shore’s working environment is flexible, collaborative, and down to earth. We work hard and deliver exceptional quality, but don’t take ourselves too seriously in doing so.

What it’s like to work at Shore:

  • Flexible culture and working environment
  • Opportunities to learn and advance
  • Contribute to innovative projects
  • Be encouraged to bring your own ideas forward
Job Description

Reporting to the Director, Platform Services, the Lead Dev Ops Engineer is a senior, hands-on technical leader responsible for building, operating, and continuously improving Afflo’s production and non-production cloud environments, CI/CD pipelines, observability, and operational toolchain. This role translates the VP’s operational strategy, standards, and compliance objectives into reliable, scalable, secure implementation while leading execution across the Dev Ops function day to day.

You will work closely with Product Engineering, QA, Service Management, Implementation/Project Delivery, and external vendors to ensure Afflo services meet uptime, performance, security, and audit expectations in regulated healthcare contexts. You will also mentor other Dev Ops engineers, lead incident response and prevention work, and drive practical improvements that reduce operational risk and accelerate safe delivery.

This role is demanding and diverse, involving:

Operational ownership of cloud infrastructure and delivery pipelines

Release engineering and environment lifecycle management

Observability, incident leadership, and continuous improvement

Security controls, evidence readiness, and DR/BCP execution

Tooling automation that reduces toil and improves team productivity

Responsibilities

Operational Ownership

Own the reliability and day-to-day operation of Afflo environments (production and non-production), ensuring uptime, performance, responsiveness, and strong operational hygiene.

Lead triage, mitigation, and restoration during incidents; coordinate with Service Management and engineering stakeholders through resolution.

Conduct and author post-incident reviews and drive prevention work to reduce recurrence, improve MTTR, and increase change safety.

Establish and maintain on-call standards, escalation paths, maintenance practices, and operational runbooks aligned with IT Operations and System Administration policies.

Design, build, and maintain secure, resilient cloud infrastructure using Infrastructure as Code (IaC) with reusable modules, review discipline, and predictable environment patterns.

Build and improve environment lifecycle workflows (provision, reset, clone, teardown) for QA/UAT/demo/customer environments and internal team needs.

Implement secure-by-default patterns: network segmentation, least privilege, secrets handling, encryption, audit logging, and access reviews.

Perform capacity planning and cost optimization—balancing availability, scalability, and operating cost, and providing actionable recommendations to the VP of Delivery.

Design, set up and maintain AI specific workloads and pipelines.

E.g. Data processing, model training, inference etc.

CI/CD, Release Engineering, and Delivery Enablement

Build and maintain automated CI/CD pipelines to enable rapid, safe deployments, including release gates, automated checks, artifact integrity, and rollback readiness.

Participate in and/or lead major release windows and maintenance deployments; ensure readiness checks, comms coordination, and post-release verification.

Standardize release processes across teams/products to reduce variance, improve predictability, and support project timelines and…

To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary