Senior Site Reliability Engineer
Listed on 2026-09-05
-
IT/Tech
SRE/Site Reliability
As a Senior Site Reliability Engineer (SRE) at CDW, you will improve the reliability, scalability, performance, security, and operational excellence of applications and platforms supporting Managed Services. This includes customer-connectivity platforms that enable CDW service delivery teams to monitor, troubleshoot, and manage customer environments.
You will serve as a senior technical escalation point, resolving complex operational issues, reducing operational demands on development teams, and partnering with software engineering, infrastructure, security, and business stakeholders to improve service resilience. This role combines software engineering, systems engineering, automation, and operational leadership to reduce risk, improve the customer experience, and accelerate delivery.
The Senior SRE will also mentor engineers, establish reliability engineering practices, and influence architectural and operational decisions across the organization.
- Drive service reliability, scalability, performance, security, and operational excellence across Managed Services applications and customer-connectivity platforms.
- Serve as a senior escalation point for complex operational issues, resolving incidents that exceed the current SRE team's expertise and reducing operational interruptions for development teams.
- Design, develop, and maintain automation using Ansible and Python to reduce operational toil, improve consistency, and increase platform reliability.
- Establish and mature reliability and observability practices, including SLIs, SLOs, Error Budgets, metrics, logs, traces, alerting, and dashboards.
- Lead major incident response, Problem Management, root cause analysis, post-incident reviews, and corrective actions that address systemic issues and prevent recurrence.
- Perform operational readiness and resiliency reviews while identifying reliability risks, operational gaps, technical debt, and opportunities for continuous improvement.
- Troubleshoot complex issues across applications, Kubernetes environments, infrastructure, networking, identity services, databases, certificates, cloud services, and third‑party integrations.
- Mentor engineers and partner with development, infrastructure, security, and business teams to improve technical standards, operational practices, documentation, release quality, and production readiness.
- Participate in a scheduled primary and secondary on‑call rotation after completing training and demonstrating readiness to independently support the environment.
- Bachelor's degree in Computer Science, Software Engineering, Information Technology, or a related field and 7+ years of experience in Software Engineering, Site Reliability Engineering, Dev Ops, Platform Engineering, or a related discipline; or 10+ years of equivalent professional experience.
- 5+ years of experience administering Linux-based systems in enterprise environments.
- 5+ years of experience developing automation solutions using Ansible.
- 3+ years of experience developing automation and operational tooling using Python.
- 3+ years of experience supporting Kubernetes or other container orchestration platforms in production environments.
- Experience supporting business‑critical production applications and distributed systems.
- Experience leading major incident response, Problem Management, root cause analysis, and corrective action initiatives.
- Experience working within ITIL‑aligned Incident, Problem, and Change Management processes.
- Experience implementing observability solutions using metrics, logs, traces, alerting, and dashboards.
- Experience defining and measuring service reliability using SLIs, SLOs, and Error Budgets.
- Strong knowledge of networking, CI/CD pipelines, source control, certificates, secrets management, and modern software delivery practices.
- Demonstrated ability to troubleshoot complex technical issues, influence technical direction, mentor engineers, and communicate effectively with technical and non‑technical stakeholders.
- Experience reviewing, troubleshooting, and making minor enhancements to existing Java‑based applications. This is not primarily a Java application‑development…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).