Lead Infrastructure Engineer
Listed on 2026-09-04
-
Software Development
Cloud Engineer - Software, DevOps, AWS
About Onos Health
Onos Health’s mission is simple but ambitious: ensure every healthcare dollar goes toward delivering the highest quality care. Today, 30% of total U.S. healthcare spending is wasted due to ineffective care and administrative burden caused by misalignment between providers and payers.
Onos is addressing this by building the largest AI-driven healthcare data platform. Our models enables payers to make faster, more accurate decisions across their populations. By guiding members to the right care, Onos is channeling more dollars to high-quality care that drives better outcomes while making healthcare more affordable.
Onos is well-funded by some of the best healthcare investors and is working with the nation’s largest health plans. Come join a category-defining company and help reimagine healthcare for the better.
Why Onos?Meaningful impact:
Help fix what is fundamentally broken in healthcareDirect collaboration:
Work alongside experienced founders with deep healthcare and data expertiseCulture:
Join a high-performing, transparent, and results-oriented teamOwnership:
Significant responsibility and autonomy from day oneOpportunity:
Play a pivotal role in building a fast-growing, category-defining healthcare AI company
We're seeking an experienced infrastructure engineer to become our first dedicated platform hire and the owner of the infrastructure the Onos platform runs on. Onos is in production with the nation's largest health plans, which comes with contractual uptime SLAs, disaster recovery commitments, and a security bar (SOC 2, HIPAA) our clients audit. Until now this has been carried collectively by our product engineers and founders — you'll own it end to end.
We build heavily with AI coding agents, so much of your leverage will come from specifying work well and directing agents rather than typing every line yourself; prior tech lead or engineering management experience translates directly. As an early team member, you'll set the patterns every future platform engineer at Onos inherits. This role is a hybrid role based in San Francisco, where you'll be expected to work at our office in person 3 times a week.
Own our availability, disaster recovery, and backup commitments to enterprise clients — multi-region failover architecture and the recovery exercises that prove it
Stand up production monitoring, alerting, SLOs, and our on-call rotation and incident response process
Own the technical controls behind SOC 2 and HIPAA: cloud security posture (AWS org guardrails, IAM least-privilege, KMS/encryption), vulnerability remediation, and continuous audit evidence through Vanta, with a path toward HITRUST
Build the CI/CD pipelines, Terraform/IaC foundations, preview environments, and test infrastructure the whole team ships on
Build the guardrails that let AI coding agents ship safely — policy-as-code, deploy verification, and agent-operated operations tooling
Set the strategy and operating rhythm for platform work: priorities, status, and what we deliberately defer
Architect AI SRE agents to ensure up-to-date compliance and reliability, enabling engineers to work more effectively and strategically
Right-size enterprise-grade reliability: SLOs and alerting you can trust without drowning a small team in pager noise
Turn compliance into continuously verified infrastructure — security controls and audit evidence as code
Scale a CI/CD and environments platform where AI agents, not just humans, are the primary users
Support multi-region disaster recovery with defined RTO/RPO targets and immutable, restore-tested backups for a multi-tenant healthcare platform
At Onos, we work with a modern tech stack where we continuously evaluate and adopt cutting-edge technologies as we scale.
Infrastructure/Systems: AWS (ECS, Bedrock, Glue, etc.), Langfuse, Terraform
Languages/Frameworks
Backend:
Python, Django, Celery / Celery Beat, django-ninja, django-tenantsFrontend:
NextJS, Typescript, Tanstack Query, Shadcn UI, Zod, Nuqs
Database/Storage: PostgreSQL (AWS RDS), S3, Clickhouse
Development Tools: Github, Linear, Claude Code, Codex,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).