Site Reliability Engineer
Listed on 2026-08-29
-
IT/Tech
SRE/Site Reliability
Senior Site Reliability Engineer
Enzo Health is a healthcare technology company transforming home health operations through purpose-built artificial intelligence. We deliver a secure, HIPAA-compliant AI platform that unifies intake, clinical documentation, coding, and quality assurance—enabling agencies to reclaim time and revenue while elevating patient care.
Enzo addresses the critical challenges facing home health agencies today: rising operational costs, clinician burnout, shrinking reimbursement margins, and increasing compliance demands. Our integrated AI solution automates documentation workflows from referral to final QA, allowing clinical staff to focus on delivering exceptional patient care.
We are hiring a Senior Site Reliability Engineer to join our Security and Site Reliability team. You will focus on a stable, scalable AWS and Kubernetes platform, reliable Postgres operations, and safer delivery. You will also help mature observability, on-call, and incident response practices across engineering.
This is a hands-on role. You will diagnose production problems, improve infrastructure, write automation, and help product engineers operate their services with confidence. You will make our systems safer without slowing product delivery.
Strengthen the production platform
- Operate and improve our AWS and Kubernetes environments.
- Own infrastructure changes through Terraform, including modules, state, review standards, and drift control.
- Improve environment management, cluster practices, and deployment reliability.
- Make CI/CD and releases safer through clear checks, reliable rollbacks, environment consistency, and release observability.
- Improve capacity planning, resource controls, resilience, and cost visibility.
- Build automation that removes repetitive operational work and reduces avoidable failures.
Improve Postgres reliability
- Own the operational health of Postgres in production.
- Verify backups and regularly test restore procedures against agreed recovery targets.
- Improve monitoring for connections, storage, slow queries, locks, and other important failure signals.
- Guide safe schema migrations, database access, and production change procedures.
- Find performance and capacity risks before they affect customers.
- Maintain clear database runbooks for common failures and emergency work.
Mature production operations
- Improve dashboards, monitors, logs, traces, and alert routing.
- Validate existing service-level indicators and objectives, then close important coverage gaps.
- Participate in and improve the company-wide on-call rotation.
- Write practical runbooks and make escalation paths clear for engineers across the company.
- Lead or support incident response, recovery, and blameless post-incident reviews.
- Use incident and reliability data to set priorities and measure improvement.
Work across Security and engineering
- Work with Security to keep infrastructure changes consistent with SOC 2 and health care security requirements.
- Apply least-privilege access, secure defaults, secrets management, encryption, and auditable change practices.
- Support vulnerability remediation, disaster-recovery exercises, and secure production access.
- Help product teams include reliability and operational risk in technical decisions.
What success looks like in the first six months
- Kubernetes, AWS, and Terraform have clear operating standards and fewer manual failure points.
- Deployments are observable, repeatable, and easy to roll back.
- Postgres backups and restores are tested, and database health and capacity risks are visible.
- Database migrations and emergency access follow safe, documented procedures.
- Production alerts are useful and actionable, with clear ownership and less noise.
- On-call responders have the runbooks, access, and escalation paths that they need.
- Production incidents result in tracked corrective work and measurable reliability gains.
Qualifications
- At least 5 years of experience in site reliability, platform, infrastructure, or production engineering.
- Strong hands-on experience with AWS and production Kubernetes.
- Strong experience with Terraform and infrastructure as code.
- Experience operating Postgres in production, including backup and…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).