Senior Manager; DevOps, Automation, GCP
Listed on 2026-07-29
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
We're building a world of health around every individual - shaping a more connected, convenient and compassionate health experience. At CVS Health®, you'll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger - helping to simplify health care one person, one family and one community at a time.
Position Summary:Join Fortune 7 CVS Health as a Sr. Manager, Software Engineering - Dev Ops, Observability & Monitoring to lead strategic initiatives for the CVS Caremark Digital team. In this role, you will drive the vision and execution of modern platform engineering capabilities that enable scalable, reliable, and secure application development. You will build and lead a high-performing engineering team responsible for designing, implementing, and operating cloud-native platforms.
The ideal candidate is a hands‑on leader with deep technical expertise in cloud architectures, Dev Ops, and observability, combined with the ability to lead transformation and operational excellence initiatives.
Leadership & People Management:
- Lead, mentor, and grow a team of software engineers and SRE/Dev Ops engineers.
- Foster a culture of accountability, innovation, and continuous improvement.
- Define team goals, OKRs, and performance metrics aligned with organizational strategy.
- Partner with product, architecture, and business stakeholders to deliver platform capabilities.
- Drive the adoption of AIOps solutions for predictive monitoring, anomaly detection, and automated root cause analysis.
- Integrate machine learning models and analytics into monitoring pipelines to proactively detect and prevent incidents.
- Develop intelligent alerting systems to reduce noise and improve signal quality.
- Architect and implement scalable observability frameworks across metrics, logs, traces, and events.
- Establish standards for instrumentation, telemetry collection, and distributed tracing.
- Enable proactive monitoring, alerting, and incident detection using tools such as Datadog, Prometheus, Grafana, and Splunk.
- Define and implement enterprise-wide SRE practices, including SLIs, SLOs, error budgets, and reliability governance.
- Drive adoption of CI/CD pipelines, Infrastructure as Code (IaC), and Git Ops practices.
- Lead the design and evolution of scalable, automated, and secure platform engineering solutions.
- Standardize development and deployment workflows across teams.
- Champion Dev Ops maturity, developer productivity, and release automation.
- Improve system reliability through error budgets, resiliency patterns, and chaos engineering.
- Lead incident response processes, postmortems, and root cause analysis.
- Drive continuous improvements in MTTR (Mean Time to Recovery) and system availability.
Infrastructure:
- Architect and manage solutions across cloud platforms (AWS, Azure, or GCP).
- Ensure scalability, security, and cost optimization of infrastructure.
- Oversee containerization and orchestration using Docker and Kubernetes.
- Integrate Dev Sec Ops practices into pipelines and platform tooling.
- Ensure compliance with enterprise security standards and regulatory requirements.
- Automate security scanning, vulnerability management, and policy enforcement.
- 5+ years of experience in software engineering, SRE, or production engineering within large‑scale distributed systems.
- Hands‑on experience with AIOps or intelligent monitoring platforms, including anomaly detection and event correlation.
- Experience with observability tools such as App Dynamics, Grafana, Prometheus, and Splunk.
- Strong expertise in cloud platforms (AWS, Azure, or GCP), cloud‑native architectures (Kubernetes, containers, microservices), and CI/CD pipelines (Git Hub Actions, Jenkins).
- Experience with Infrastructure as Code (Terraform, AWS Cloud Formation, or GCP Deployment Manager).
- Proficiency in at least one programming language (e.g., Python, Java, Go).
- Proven track record…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).