Software Engineer - SRE, Retail & Pharmacy
Listed on 2026-08-22
-
IT/Tech
SRE/Site Reliability, Systems Engineer
We're building a world of health around every individual - shaping a more connected, convenient and compassionate health experience. At CVS Health®, you'll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger - helping to simplify health care one person, one family and one community at a time.
Position Summary:
About the Team
Our SRE team is the execution engine behind the reliability, availability, and performance of distributed store technology powering thousands of retail and pharmacy locations nationwide. We operate across pharmacy platforms, Point of Sale (POS) systems, handheld devices, store servers, dispensing systems, and edge computing infrastructure in hybrid cloud and on-premises environments deployed at fleet scale.
Our engineering philosophy is grounded in five pillars:
Detection
, Prevention
, Recovery
, Learning Loops
, and Developer Experience (DevX).
Our operating principle is what we call the reliability covenant
: our success is not measured by how many incidents we respond to, it is measured by how much reliability capability we transfer to the engineering teams we serve. The goal is development teams that carry reliability ownership independently, not teams that rely on SRE to keep their services running. If you are drawn to building capability that outlasts your direct involvement, this team is built for that purpose.
We track operational toil as an engineering metric
, not as a permanent operational reality. Engineers are expected to identify recurring manual work, eliminate it through automation, and document the reduction. Toil accumulation is treated as a reliability risk and a capacity cost.
About the Role
As a Software Engineer - SRE
, you are a practitioner-level contributor focused on building and running reliable distributed systems. You work within assigned services and domains implementing observability, improving alerting quality, responding to incidents, writing automation, and contributing to the reliability programs that run across the organization.
Scope: Service and task level, you execute with direction and grow toward autonomous ownership.
The environment you are joining
This role exists inside an active SRE transformation. Many of the systems you will monitor, the processes you will contribute to, and the tool chains you will use are being built or significantly improved in parallel with the day-to-day operational work. You will contribute to defining processes as much as following them. Comfort with ambiguity and a bias toward building, not just operating is essential to success in this role.
Engineers who thrive here find that environment energizing, not frustrating.
The operating environment includes an edge computing fleet deployed directly inside store locations, unattended nodes that cannot be reached by on-site SRE engineers. This means a deployment or configuration change that goes wrong can simultaneously affect thousands of locations. You will develop a fleet operations mindset alongside a service reliability mindset: blast radius is geographic, not just functional.
What You Will Do
Detection & Observability
- Implement alerting and dashboards using Prometheus
, Grafana
, Loki
, Jaeger
, and Open Telemetry to provide visibility into distributed retail and pharmacy systems - Contribute to Service Level Indicator (SLI) and Service Level Objective (SLO) definition for assigned services under Senior(s) guidance; learn error budget mechanics and how burn rate translates to patient and customer impact
- Build and maintain service dashboards covering the golden signals: availability, latency, error rate, and saturation anchored to Critical User Journey (CUJ) outcomes, not just infrastructure metrics
- Participate in alert tuning exercises; document false-positive patterns and noise sources with enough specificity to enable SSE-level remediation
- Learn the principles of anomaly-based detection: understand the difference between threshold alerting and time-series baseline deviation, and how ML-generated signals differ from rule-based alerts
Prevention & Reliability Engineering
- Participate in Production Readiness Reviews (PRR); execute assigned checklist items, contribute findings, and understand the rationale behind each gate
- Write unit and integration tests for SRE tooling and automation scripts at production quality
- Follow and actively improve existing operational runbooks; flag gaps, missing failure modes, outdated steps, and ambiguous procedures for remediation, don't just identify, propose the fix
- Support performance testing and reliability audits under SSE direction; develop hands‑on methodology exposure alongside execution skills
- Apply standard reliability engineering patterns: circuit breakers, retries with exponential backoff, timeouts, and bulkhead isolation etc in code and configuration contributions
- Build awareness of fleet-scale deployment risk:…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).