Site Reliability Engineer (SRE) – Production Services
Listed on 2026-08-05
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Automation & Efficiency
Automate the top 5 high-volume support and request types
Build self-service and agent-driven solutions to reduce manual work
Harden operational workflows for consistency, auditability, and resilience
Implement auto-retry and backoff for recurring failure patterns
Reliability EngineeringDefine and manage Service Level Objectives (SLOs) for critical services and batch processes
Apply error budget concepts to guide reliability and release decisions
Improve batch reliability through standardized recovery patterns and monitoring
Observability & MetricsBuild reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage
Improve operational reporting and visibility across incidents, problems, and changes
Runbooks & Self-ServiceDevelop and expand runbooks for key production scenarios
Convert runbooks into automated remediation workflows
Enable self-service for repeat operational requests
Drive conversion of repeat incidents into permanent fixes and known problems
Self-Healing & Intelligent OperationsImplement self-healing capabilities to minimize manual intervention
Optimize alerting systems (e.g., Moogsoft) to reduce noise and improve signal quality
Leverage automation and AI to resolve recurring issues with minimal human involvement
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).