×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer Remote

Remote / Online - Candidates ideally in
Toronto, Ontario, C6A, Canada
Listing for: AXON-Networks
Remote/Work from Home position
Listed on 2026-10-02
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Network Engineer
Salary/Wage Range or Industry Benchmark: 110000 - 152000 CAD Yearly CAD 110000.00 152000.00 YEAR
Job Description & How to Apply Below
Position: Site Reliability Engineer Remote)

AXON Networks delivers a robust AI-driven, analytics-based orchestration platform and a wide portfolio of next-gen high-speed routers that leverage the newest Wi‑Fi technologies. Together, these technologies give ISPs the ability to manage and troubleshoot their networks in real time, and to deliver an outstanding customer experience.

AXON Networks is a trusted strategic partner for its customers, helping them evaluate their current technologies and business models, and creating and executing strategies that enable them to innovate faster, accelerate their digital transformations, and strengthen their relationships with consumers.

AXON Networks is headquartered in Irvine, CA USA with Asia HQ in Singapore and also operating in Denmark, Spain and Vietnam.

The Site Reliability Engineer will improve the availability, performance, scalability and recoverability of AXON Networks cloud solutions. You will combine software engineering with hands‑on NOC operations to make the complete cloud‑to‑device service path observable, supportable and resilient at fleet scale.

You will help establish practical SRE capabilities inside the NOC while partnering closely with Support, Operations, cloud and Dev Ops Engineering. You will participate in a sustainable on‑call rotation and improve the NOC’s ability to diagnose customer‑impacting issues.

Role mandate
  • Own reliability outcomes for assigned cloud services
  • Improve observability, capacity, resilience and recovery
  • Define and operationalize service‑level indicators, service‑level objectives and actionable alerting.
  • Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations and routine production changes.
  • Lead technically during incidents, drive evidence‑based learning and ensure high‑value corrective actions are completed.
What you will own
  • Establish reliability baselines, SLIs, SLOs and error budgets for cloud services and critical device‑management workflows such as onboarding, provisioning, configuration, telemetry collection, command execution and firmware delivery.
  • Trace failures across the end‑to‑end service path: cloud APIs and microservices, Kubernetes and infrastructure, databases and messaging, internet and access‑network dependencies, device‑management protocols and the devices
  • Identify fleet‑wide and customer‑specific failure patterns involving device reachability, session stability, configuration drift, command latency, telemetry gaps, firmware behaviour and cloud capacity.
  • Contribute operability requirements and production evidence during design and readiness reviews
  • Maintain NOC dashboards for service health, device reachability, provisioning success, command and telemetry performance, firmware adoption and customer impact.
  • Participate in the NOC production on‑call rotation and serve as a technical incident lead or senior troubleshooter when appropriate.
  • Diagnose complex failures across applications, cloud infrastructure, Kubernetes, APIs, networking, DNS/TLS, databases, messaging platforms, device‑management sessions and CPE behavior.
  • Coordinate evidence gathering and technical escalation with service‑provider customers, Engineering, firmware, Dev Ops and vendors while maintaining clear mitigation, recovery and handoff.
  • Lead or contribute to post‑incident reviews; convert recurring device, platform and process failures into prioritized and measurable corrective actions.
  • Develop production‑grade software, scripts and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance and recovery.
  • Improve CI/CD and Git Ops practices for operational software and infrastructure, including automated testing, release validation, progressive delivery and rollback readiness.
  • Manage or…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary