×
Register Here to Apply for Jobs or Post Jobs. X

SRE​/DevOps Engineer - 68330

Job in Toronto, Ontario, C6A, Canada
Listing for: Hitachi Digital Services
Full Time position
Listed on 2026-08-04
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Support
Salary/Wage Range or Industry Benchmark: 90000 - 140000 CAD Yearly CAD 90000.00 140000.00 YEAR
Job Description & How to Apply Below

Our Company

We're Hitachi Digital Services, a global digital solutions and transformation business with a bold vision of our world's potential. We're people-centric and here to power good. Every day, we future-proof urban spaces, conserve natural resources, protect rainforests, and save lives. This is a world where innovation, technology, and deep expertise come together to take our company and customers from what's now to what's next.

We make it happen through the power of acceleration.

Imagine the sheer breadth of talent it takes to bring a better tomorrow closer to today. We don't expect you to 'fit' every requirement – your life experience, character, perspective, and passion for achieving great things in the world are equally as important to us.

Job Description

Meet Our Team Join our Site Reliability Engineering (SRE) Operations team, where reliability, automation, and operational excellence are at the heart of everything we do. We ensure the stability, availability, and performance of enterprise applications running across modern cloud‑native and hybrid platforms, including Kubernetes, APIs, cloud services, databases, Kafka, and API gateways.

As an L1 SRE Operations Engineer
, you'll be the first line of defense, monitoring production environments, responding to alerts, executing operational runbooks, and partnering with senior engineers to maintain highly available and resilient platforms. This is an excellent opportunity for professionals looking to build hands‑on experience in cloud operations, Dev Ops, and Site Reliability Engineering.

What You'll Be Doing
  • Monitor enterprise applications, infrastructure, dashboards, logs, and alerts across cloud and on‑premises environments.
  • Perform first‑level incident triage by analyzing alerts, collecting logs and metrics, and determining whether issues are application or platform related.
  • Execute standardized operational runbooks for incident resolution, deployments, maintenance activities, and routine operational tasks.
  • Monitor and support Kubernetes environments by validating pod health, deployments, name spaces, logs, and service endpoints.
  • Troubleshoot infrastructure and application issues using Linux utilities, networking tools, and monitoring platforms.
  • Escalate complex incidents to L2/L3 engineering teams with complete diagnostic information to accelerate resolution.
  • Support API gateways, web application firewalls (WAF), Kafka platforms, databases, and cloud infrastructure across AWS, Azure, and GCP.
  • Maintain accurate incident documentation, operational records, and knowledge base updates while identifying opportunities to improve runbooks and automation.
  • Collaborate with development, platform engineering, and infrastructure teams during incident response and production support.
  • Assist with onboarding new applications into the operational support framework while ensuring monitoring, alerting, and operational readiness.
  • Contribute to continuous improvement by identifying repetitive manual activities suitable for automation.
  • Provide timely and professional communication to stakeholders during production incidents and operational events.
What You'll Bring to the Team

Required Qualifications
  • 2–5 years of experience in IT Operations, NOC, SRE, Dev Ops, or Infrastructure Support.
  • Working knowledge of Kubernetes administration and day‑to‑day cluster operations.
  • Good understanding of Linux administration and command‑line troubleshooting.
  • Familiarity with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
  • Experience with observability and monitoring tools such as Prometheus, Grafana, Splunk, ELK Stack, Datadog, Argos, or AIOps platforms.
  • Ability to execute operational runbooks and follow structured incident response procedures.
  • Experience using Kubernetes CLI (kubectl) to verify pod health, deployments, name spaces, and application logs.
  • Basic scripting knowledge in Python, Bash, or Power Shell for operational automation.
  • Understanding of networking fundamentals including DNS, HTTP/HTTPS, TCP/IP, firewalls, WAF, proxies, connectivity troubleshooting, and diagnostic tools such as ping, curl, netstat, and trace route.
  • Strong analytical and troubleshooting…
Note that applications are not being accepted from your jurisdiction for this job currently via this jobsite. Candidate preferences are the decision of the Employer or Recruiting Agent, and are controlled by them alone.
To Search, View & Apply for jobs on this site that accept applications from your location or country, tap here to make a Search:
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary