Site Reliability Engineer
Listed on 2026-07-23
-
IT/Tech
SRE/Site Reliability, IT Support, Cloud Computing: Infrastructure & Operations, Systems Engineer
Positoin:
Site Reliability Engineer Onsite in Blue Ash, OH Long-term Contract W2 Pay: $60-$68/hr Skills
- 3+ years of experience in Site Reliability Engineering (SRE), Production Support, Incident Management, or Systems Operations
- Hands‑on experience leading major incidents (P1/P2 outages), including bridge/war‑room facilitation, stakeholder communication, and incident coordination
- Experience conducting Root Cause Analysis (RCA) and driving corrective and preventative actions through completion
- Strong understanding of monitoring and observability tools such as Dynatrace, Azure Monitor, Datadog, Grafana, New Relic, or similar platforms
- Experience troubleshooting production issues across cloud, on‑premises, and distributed systems environments
- Working knowledge of Linux administration and scripting (Bash, Python, or similar)
- Experience working with Kubernetes and Docker in production environments
We are seeking a Site Reliability Engineer (SRE) focused on Major Incident Management, Problem Management, and Production Reliability across cloud, on‑premises, and retail store environments. This individual will serve as the SRE lead during high‑severity incidents, drive root cause analysis efforts, partner with engineering teams to improve platform resilience, and help reduce the recurrence of production issues. This role requires a strong troubleshooting mindset, experience coordinating incident response across multiple teams, and the ability to communicate effectively with both technical and business stakeholders.
Key Responsibilities- Lead and coordinate response efforts during major production incidents, serving as the primary SRE representative on incident bridges
- Drive incident triage, escalation, communication, and resolution activities across multiple technology teams
- Conduct root cause investigations and facilitate Problem Management processes to eliminate recurring issues
- Partner with engineering teams to identify systemic reliability gaps and implement long‑term improvements
- Collaborate with stakeholders to define, measure, and improve service reliability using SLOs and SLIs
Exact compensation may vary based on several factors, including skills, experience, and education. Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).