Site Reliability Engineer; SRE
Listed on 2026-07-09
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support
Job Title
Responsibilities:
Develop, test, and debug automated tasks (Apps, Systems, Infrastructure)
Troubleshoot minor incidents and contribute to resolution through post-mortems
Participate in the application or service development lifecycle through code contributions
Engage with tools and operations teams to address failure patterns and incidents
Develop automation tools for efficient, noiseless alerting, toil, and technical debt
Conduct performance tests, document and/or identify application optimizations
Qualifications:
Bachelor’s degree or equivalent experience in an software engineering discipline
Proficiency in at least one software language (e.g. Java, Python, etc.)
Proficiency in SQL – preferably with Oracle
Strong problem solving abilities
Understanding of the software delivery lifecycle
Expertise in application, data, and infrastructure architecture disciplines
Knowledge of one or more infrastructure components (e.g. networking, cloud services, orchestration tools, containerization, compute, and storage systems)
Capable of managing service-level changes to a system or service
Hands-on experience with cloud deployment, monitoring, and ops analysis tools such as Kubernetes, Prometheus, Elasticsearch, Grafana, Kibana, Splunk, Dyna Trace, etc.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).