SRE Leader
Listed on 2026-07-24
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
We are makers and builders.
We challenge the status quo and have fun doing it.
We bring passion and purpose to making a difference in healthcare, elevating how care is delivered and experienced.
Every day, intelligent software orchestrates the physical world around us – from matching drivers with riders to optimizing global supply chains. Yet inside hospitals, where every second and every decision can affect a patient's life, operations are still spread across dozens of disconnected systems.
, we're changing that. By combining proprietary hardware, AI-powered intelligence, and deep integrations with the systems hospitals already rely on, we're creating a real-time understanding of hospital operations that software alone can't deliver. That intelligence powers the execution layer hospitals have been missing – helping care teams make smarter decisions and deliver better patient care.
Backed by Goldman Sachs and trusted by leading health systems including HCA Healthcare, Sutter Health, Advent Health, Trinity Health, Northwell Health, Cleveland Clinic, and the U.S. Department of Veterans Affairs, we're pioneering the next generation of healthcare operations. We've more than doubled our revenue over the past year and are on track to surpass $70M in annual recurring revenue – not because we're following a market, but because we're defining one.
If you're excited to solve hard problems, work with a team of builders, and help hospitals deliver better care every day, we'd love to meet you!
We’re looking for an SRE Leader to own the reliability, performance, and automation of our cloud‑based, real‑time platform. This role will focus on keeping our platform running smoothly 24/7, minimizing downtime, improving observability, incident response, and self‑healing automation. You will lead and scale the SRE team to ensure our infrastructure stays ahead of demand, operates efficiently, and meets the needs of our growing healthcare customers.
Responsibilities- Ensure 99.99% uptime across our cloud platform, meeting strict SLAs for healthcare customers.
- Design and implement self‑healing, fault‑tolerant systems to prevent failures before they happen.
- Define SLIs, SLOs, and SLAs
, ensuring proactive performance monitoring and incident resolution. - Architect and manage scalable cloud infrastructure (AWS) for massive real‑time data processing.
- Optimize containerized environments (Kubernetes, Docker) to support multi‑region deployments.
- Lead the adoption of infrastructure as code (Terraform) to fully automate infrastructure management.
- Build and refine a world‑class monitoring, alerting, and logging system using Prometheus, Grafana, Open Telemetry, and Datadog.
- Lead incident response and on‑call operations
, reducing mean time to detection (MTTD) and mean time to resolution (MTTR). - Conduct blameless postmortems and continuously improve system resilience.
- Reduce manual intervention through automated deployment, scaling, and failover mechanisms.
- Partner with Security & Compliance teams to ensure infrastructure meets HIPAA and SOC 2 standards.
- Lead disaster recovery and business continuity planning to ensure critical healthcare services are always available.
- Drive technical strategy and roadmap for scalability, monitoring, and reliability engineering.
- Collaborate with Product, Engineering, and Infrastructure teams to align SRE initiatives with business priorities.
- 10+ years of experience in Site Reliability Engineering or Cloud Infrastructure.
- Proven success scaling high‑traffic, mission‑critical platforms in SaaS, IoT, or healthcare.
- Deep expertise in cloud platforms (AWS), Kubernetes, and distributed systems.
- Strong background in monitoring, logging, and observability with Prometheus, Open Telemetry, or similar tools.
- Hands‑on experience with incident management, postmortems, and building resilient systems.
- Deep knowledge of CI/CD automation, Git Ops, and infrastructure as code (Terraform, etc.).
- A mature leadership approach
, with the ability to drive technical strategy while growing and mentoring a high‑performance SRE team. - Strong understanding of network security, access management, and compliance frameworks (HIPAA, SOC
2).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).