Acquire-Site Reliability Engineer
Listed on 2026-07-16
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Acquire-Site Reliability Engineer
Salary Range: $90,000.00 To $ Annually
Site Reliability EngineerAcquire Learning is a learning management platform built for ABA (Applied Behavior Analysis) therapy. Clinicians and behavior technicians use it daily with clients on the autism spectrum, and the data it captures shapes real treatment decisions. We are a small, product-focused team in a HIPAA-regulated environment. When a clinician is mid-session with a child, the platform behaving predictably is the difference between productive therapy and a disrupted session.
Your work affects the communities we serve.
Acquire is hiring its first dedicated Site Reliability Engineer, a mid-level role with a clear path to Lead SRE as we grow. You will own production health, a trustworthy release pipeline, and the reliability surface of the codebase, while raising release-quality risk and acting as the customer-facing escalation point. You report to the Lead Engineer and work regularly with the CEO and CTO.
You will not inherit a mature SRE team or a thick runbook library; you will help build them.
Just as important as the technical background: we want a product-focused thinker. AI tooling makes raw implementation cheaper, so the scarce skill is judgment about why the product matters to clinicians and what actually needs building. Treat reliability and ops as ways to keep a great product healthy, and step into feature work when the team needs it.
This is a broad role today by design. As the team grows it narrows toward Lead SRE: reliability strategy, incident response, and how ops, observability, and release engineering work at Acquire.
What You'll DoProduction reliability: own day-to-day health across AWS and MongoDB Atlas. Triage and respond to alerts (Cloud Watch, Sentry, Google Chat ops-alerts), run root-cause and incident comms, and turn retros into runbooks and alerting improvements. Close HIPAA-aware observability gaps: PHI-safe logging, auditability, access controls, incident evidence.
Dev Ops and release engineering: operate and improve our Git Hub Actions deploy pipelines (backend, webapp, native), maintain Terraform infrastructure and review infra PRs for safety, improve CI signal quality, coordinate mobile releases via Test Flight and Google Play, and harden rollback, restore, and break-glass paths.
Reliability-focused code (Type Script/Node): repair scripts, migrations and index management, observability instrumentation, tenant-scoped operational tooling, background job lifecycle, deploy tooling, and e2e (Playwright) and integration tests. This is your primary lane, but with agentic tooling we expect you to help on product features when it counts rather than treating "that's not infra" as a boundary.
Release quality and QA: validate release candidates, walk core clinical workflows on web, iOS, and Android, run our test suites (Playwright, Jest, Postman), and drive release checklists and post-release verification.
Customer escalations: be the first internal contact for customer-reported issues. Reproduce, isolate, document, prioritize by clinical impact, and own the loop back to the customer.
Where We Need HelpHIPAA-aware operations, multi-tenant architecture (tenant isolation, org-safe migrations and diagnostics), disaster recovery and restore confidence, data-integrity operations, security and production-access hygiene (IAM, secrets, least privilege), incident-response maturity (severity levels, SLOs, alerting standards), and scale and cost visibility across AWS and MongoDB.
Who You Are3-5 years that meaningfully includes SRE, Dev Ops, production, platform, or infrastructure engineering.
AWS proficiency is a hard requirement. Real, hands-on experience operating production workloads on AWS, and you can talk through it in specifics.
A product-focused thinker who asks why the product matters and what needs building, not only how the infrastructure runs, and who can contribute to feature work with agentic tooling when needed.
Self-sufficient and a fast ramper. Dropped into an unfamiliar codebase, cloud account, or toolchain, you have a real strategy to get productive on your own, and you should expect to demonstrate it live during interviews.
Curious about ABA, our product, and AI. Genuine interest in the clinical work Acquire supports is preferred, and we want people who are eager to learn the domain and to explore how AI can move it forward.
Honest about your background, with recent, checkable professional references.
Comfortable with CI/CD, Terraform or comparable IaC, MongoDB or another production database, and observability tooling (Cloud Watch, Datadog, Sentry, Grafana).
Fluent enough in Node/Type Script for reliability work: repair scripts, migrations, instrumentation, job lifecycle, deploy tooling, and e2e/integration tests.
At home on the command line, in production logs, and in cloud consoles, with real incident experience you can talk through.
Careful about production access, customer data, and tenant boundaries,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).