×
Register Here to Apply for Jobs or Post Jobs. X

Sr. Specialist, SRE – Compute Platforms

Job in Toronto, Ontario, M5A, Canada
Listing for: Canadian Tire
Full Time position
Listed on 2026-09-02
Job specializations:
  • IT/Tech
    SRE/Site Reliability, IT Infrastructure, Systems Engineer, Disaster Recovery IT
Job Description & How to Apply Below
Toronto, ON
Full time
JR
What You'll Do:
The Sr. Specialist, SRE – Compute Platforms serves as the enterprise technical owner for CTC's compute platforms, including IBM Mainframe (z/OS), AIX, IBM iSeries, IBM NOI, New Relic, and associated business-critical services and applications.
This role is accountable for platform governance, service ownership, lifecycle strategy, observability governance, vendor oversight, and reliability outcomes across the compute environment. Acting as the internal authority for compute platforms, the position ensures services remain reliable, resilient, secure, observable, supportable, and aligned with enterprise technology strategy, risk management, and modernization objectives.
Operational execution, administration, monitoring response, maintenance activities, and day-to-day infrastructure support are performed by HCL/Harmony and other designated support providers. The Sr. Specialist provides governance, strategic direction, technical leadership, and vendor accountability to ensure services are delivered in accordance with enterprise standards, contractual commitments, and business requirements.
The role leads platform roadmap development, technology currency initiatives, service reliability improvements, observability strategy, operational governance, vendor management, and continuous improvement programs while promoting Site Reliability Engineering (SRE) principles, operational excellence, automation, and platform sustainability. The position drives platform maturity, reduces consultant and key-person dependency, and establishes clear ownership and accountability across the enterprise compute environment.
Mainframe Platform Ownership and Lifecycle Governance
Provide governance and oversight of HCL/Harmony-delivered Mainframe services, including z/OS lifecycle planning, platform currency, capacity strategy, LPAR governance, and infrastructure roadmap planning.
Assess current platform, operating system, middleware, and Mainframe-supported application versions; identify lifecycle, supportability, operational, security, and compliance risks; and develop renewal and modernization strategies.
Lead platform lifecycle governance activities, including hardware refresh planning, operating system upgrade strategy, maintenance governance, change oversight, and long-term roadmap alignment.
Govern service ownership, support models, operational accountability, and vendor-delivered support services for Mainframe-hosted and Mainframe-dependent applications.
Define, govern, and periodically review monitoring and observability requirements for Mainframe infrastructure, middleware, and business-critical services.
Provide governance and oversight of disaster recovery readiness, resiliency planning, recovery testing, and recovery capability validation delivered by managed service providers.
Identify and drive resolution of operational ownership, support model, documentation, and Statement of Work (SOW) gaps through the appropriate vendors, support teams, and stakeholders.
Provide technical leadership and governance during major incidents, problem investigations, and corrective action planning, ensuring responsible teams execute required remediation activities.
Serve as the primary technical escalation and governance authority for Mainframe-related risks, service concerns, lifecycle issues, and vendor performance matters.
Monitoring, Event Management and Observability Governance
Provide governance and strategic oversight of enterprise monitoring, event management, and observability capabilities.
Establish governance standards for monitoring coverage, alerting requirements, event correlation, escalation models, service health dashboards, and operational reporting.
Drive continuous improvement of observability maturity, service visibility, monitoring effectiveness, synthetic monitoring capabilities, and operational intelligence.
Govern event management practices, including alert quality, escalation effectiveness, incident correlation, operational readiness, and service monitoring standards.
Define strategic direction and adoption roadmaps for observability platforms, monitoring technologies, event management tooling, and automation capabilities.
Act as the Compute SRE representative for enterprise monitoring strategy, operational intelligence, and observability initiatives.
Managed Service Governance and Platform Accountability
Provide technical governance, vendor oversight, and escalation leadership for HCL/Harmony-managed services across Mainframe, AIX, IBM iSeries, storage, backup, monitoring, and supporting infrastructure platforms.
Validate vendor-delivered services against contractual obligations, SOW commitments, service level expectations, operational controls, monitoring standards, and governance requirements.
Lead vendor performance reviews, operational scorecards, service reporting reviews, incident follow-ups, and continuous improvement initiatives.
Identify ownership gaps, operational risks, monitoring…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary