×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer

Job in Markham, Ontario, I3P, Canada
Listing for: Redwood Software
Full Time position
Listed on 2026-10-01
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Salary/Wage Range or Industry Benchmark: 125000 CAD Yearly CAD 125000.00 YEAR
Job Description & How to Apply Below
OUR MISSION
At Redwood, we empower our customers with lights-out automation for their mission-critical business processes.

About Us
Redwood Software is the leading orchestration platform for the autonomous enterprise, driving business transformation at the lowest total cost of ownership. Redwood empowers organizations to intelligently automate and orchestrate mission-critical business and IT processes across complex ERP, hybrid cloud, data and emerging agentic AI systems. Through its SaaS-first automation fabric—with AI embedded across the automation lifecycle—Redwood accelerates the path to autonomous operations.

Backed by 30 years of experience and trusted by more than 50% of the Fortune 50, Redwood helps organizations unlock human potential to focus on innovation, growth and what’s next.

CORE VALUES

One Team. One Redwood

Make Your Own Weather

Obsess over Customer Success

Work the Problem

Be Curious

Own the Outcome

Respect Each Other

YOUR IMPACT

Provide day-to-day management of system alerts, monitor system health, and elevate issues as necessary to maintain high availability.

Participate in a 24x7, team-shared on‑call rotation for critical SaaS platform incidents and provide support during emergencies.

Lead incident response efforts to ensure fast and effective mitigation and resolution of production issues.

Perform thorough Root Cause Analysis (RCA) and lead blameless post‑mortems to identify systemic weaknesses and establish corrective actions that prevent recurrence.

Collaborate with engineering teams to establish and enforce error budgets derived from Service Level Objectives (SLOs), balancing development velocity with system stability.

Automate routine operational tasks to reduce manual effort and toil while increasing team efficiency.

Design, deploy, and maintain cloud infrastructure using Infrastructure as Code (IaC), leveraging Terraform and Helm for deployment to EKS/Kubernetes clusters.

Design, secure, and troubleshoot AWS cloud network architecture, including VPCs, subnetting, routing, security groups/NACLs, load balancers (ALB/NLB), and VPN/Transit Gateway connectivity across a multi‑region, multi‑account environment supporting EKS and Docker Swarm on EC2.

Improve infrastructure health by developing and implementing checks, scripts, and automated remediation to proactively address known issues and enable platform self‑healing.

Maintain, develop, and evolve Continuous Integration/Continuous Delivery (CI/CD) deployment code and pipelines.

Maintain existing infrastructure running on Docker and Docker Swarm while contributing to migration strategies toward EKS/Kubernetes.

Implement and integrate new technologies and services into Redwood’s cloud infrastructure to enhance platform capabilities and resilience.

Design and implement comprehensive observability strategies across metrics, logs, and traces.

Create and refine robust monitoring and alerting configurations within the EKS/Kubernetes ecosystem.

Utilize and maintain Datadog to gather performance data, create synthetic tests, configure monitoring, and visualize system health through dashboards.

Leverage existing monitoring solutions, including Grafana and Prometheus, while supporting the migration or integration of data into a unified observability platform.

Document issues, remediation steps, system architecture, and runbooks to support knowledge transfer and rapid incident response.

Collaborate closely with Support, Customer Success, Migration, and Professional Services teams to deliver a high level of SaaS service and minimize customer impact during changes.

Maintain a strong customer focus when planning deployments and updates, considering the impact on end users before implementing changes.

Your Experience

5+ years of experience in Site…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary