SRE
Listed on 2026-07-01
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Job Title
Responsibilities: 5-7 years of experience. Gathers and analyzes metrics from monitoring platforms to assist in performance tuning and fault tolerance. Participates in system design, platform management and capacity planning. Balances feature development speed and reliability with service-level objectives. Works closely with the incident response team and restoring service to normal operation. Understands debugging and applying troubleshooting skills. Investigates, blocks and rate-limits unwanted traffic.
Utilizes monitoring systems and dashboards for proactive changes and alerting. Establishes continuous process improvement cycles where the process, performance, and supporting technologies are reviewed and enhanced where applicable. Partners with development teams to improve services through testing and release procedures.
Knowledge, Skills, Abilities:
Understanding of Kubernetes, containers, clusters and elastic scalability. Expertise in SRE principles. Mindset of continually finding ways to drive scalability, stability and performance. Cloud Services experience with Google Cloud Platform (GCP). Experience with API, service-based or microservice-based architecture. Proficiency in infrastructure, network, database, operating systems or security troubleshooting and remediation. Architecture-level knowledge of Windows and Linux and Infrastructure systems. Experience with production deployment, monitoring and operational support for enterprise-class applications (Dynatrace a plus).
Experience working with Continuous Integration/ Continuous Deployment tools. Experience in performance diagnostics, capacity planning, performance architecture design, performance tuning and performance monitoring.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).