Acquire-Site Reliability Engineer
Listed on 2026-07-24
-
IT/Tech
SRE/Site Reliability, AWS
Acquire Site Reliability Engineer
Location:
Denver, CO
Acquire Learning is a learning management platform built for ABA (Applied Behavior Analysis) therapy. Clinicians and behavior technicians use it daily with clients on the autism spectrum, and the data it captures shapes real treatment decisions. We are a small, product‑focused team in a HIPAA‑regulated environment.
About the RoleAcquire is hiring its first dedicated Site Reliability Engineer, a mid‑level role with a clear path to Lead SRE as we grow. You will own production health, a trustworthy release pipeline, and the reliability surface of the codebase, while raising release‑quality risk and acting as the customer‑facing escalation point. You report to the Lead Engineer and work regularly with the CEO and CTO.
You will help build an SRE team and comprehensive runbooks.
- Own day‑to‑day production reliability across AWS and MongoDB Atlas; triage and respond to alerts, run root‑cause analysis, incident communications, and turn retros into runbooks and alerting improvements. Close HIPAA‑aware observability gaps: PHI‑safe logging, auditability, access controls, incident evidence.
- Operate and improve our Git Hub Actions deploy pipelines (backend, webapp, native); maintain Terraform infrastructure; review infra PRs for safety; improve CI signal quality; coordinate mobile releases via Test Flight and Google Play; harden rollback, restore, and break‑glass paths.
- Repair reliability‑focused code in Type Script/Node: scripts, migrations, index management, observability instrumentation, tenant‑scoped tooling, background job lifecycle, deploy tooling, and e2e (Playwright) and integration tests. Collaborate on product features when needed.
- Validate release candidates, walk core clinical workflows on web, iOS, and Android, run our test suites (Playwright, Jest, Postman), and drive release checklists and post‑release verification.
- Serve as the first internal contact for customer‑reported issues: reproduce, isolate, document, prioritize by clinical impact, and own the loop back to the customer.
- HIPAA‑aware operations, multi‑tenant architecture (tenant isolation, org‑safe migrations and diagnostics), disaster recovery and restore confidence, data‑integrity operations, security and production‑access hygiene (IAM, secrets, least privilege), incident‑response maturity (severity levels, SLOs, alerting standards), and scale and cost visibility across AWS and MongoDB.
- 3–5 years of meaningful SRE, Dev Ops, production, platform, or infrastructure engineering experience.
- Strong proficiency with AWS; hands‑on experience operating production workloads on AWS, and ability to discuss specifics.
- A product‑focused thinker who asks why the product matters and what needs building, not just how infrastructure runs, and can contribute to feature work with agentic tooling.
- Self‑sufficient and a fast ramper; can get productive with minimal guidance and demonstrate strategy in interviews.
- Curious about ABA, our product, and AI. Genuine interest in the clinical work Acquire supports is preferred.
- Honest about your background with recent, checkable professional references.
- Comfortable with CI/CD, Terraform or comparable IaC, MongoDB or another production database, and observability tooling (Cloud Watch, Datadog, Sentry, Grafana).
- Fluent in Node/Type Script for reliability work: scripts, migrations, instrumentation, job lifecycle, deploy tooling, and e2e/integration tests.
- Effective in production logs, cloud consoles, and real incident experiences.
- Attentive to production access, customer data, and tenant boundaries, with clear understanding of diagnostic versus data‑leak risks.
- Comfortable on a small team where some process exists, some needs creating, and everyone stays close to the product.
We use AI where it genuinely helps. You do not need to be an AI expert, and we are not looking for someone who treats AI as a substitute for judgment. We want someone comfortable with tools like Claude or Cursor to expedite log triage, draft repair scripts and tests, investigate support cases, verify checklists, and work past the edge of their expertise…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).