Senior Site Reliability Engineer
Listed on 2026-09-07
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations, Cybersecurity
About NOCD
NOCD is the #1 telehealth provider for the treatment of obsessive-compulsive disorder (OCD). OCD is one of the most severe, prevalent, and misunderstood mental health conditions. NOCD creates access to online therapy for people with OCD through our telehealth platform. In the NOCD app, Members can quickly access and schedule live, face-to-face video therapy sessions with our national network of licensed Therapists that specialize in Exposure and Response Prevention Therapy (ERP) - considered the "gold standard" in OCD treatment.
At NOCD, we help people reclaim their lives with clinically proven OCD treatment, by removing barriers to OCD care, and reducing the stigma associated with OCD. We’re changing the world and need other like-minded individuals to accelerate and expand our efforts.
Chicago, IL (Hybrid 3X a week)
Senior Site Reliability Engineer (SRE)Chicago, IL (Hybrid)
Opportunity OverviewNOCD is looking for a Senior Site Reliability Engineer (SRE) to help build and scale the technology infrastructure powering our digital behavioral health platform.
This is a highly hands‑on engineering role at the intersection of software engineering, cloud infrastructure, platform engineering, reliability, and security
. You’ll help design and operate systems that are secure, scalable, observable, and resilient while enabling our engineering teams to ship software faster and more reliably.
You’ll have significant ownership in shaping our architecture, engineering standards, infrastructure, and developer experience as NOCD continues to scale.
What You'll DoSoftware Engineering & Systems
- Design, build, and maintain reliable production systems and services.
- Write production-quality code in Python, Type Script, or similar languages
. - Develop and maintain APIs, microservices, automation, and internal engineering tools.
- Contribute to architecture decisions, technical design reviews, code reviews, and engineering standards.
- Apply software engineering principles to infrastructure, automation, and reliability challenges.
Cloud & Platform Engineering
- Design, build, and operate AWS infrastructure supporting production applications.
- Manage infrastructure as code using Terraform
. - Build and maintain containerized environments using Docker and Kubernetes
. - Design and improve CI/CD pipelines, deployment automation, and release processes.
- Build internal tooling and platform capabilities that improve developer productivity.
- Help establish scalable infrastructure patterns that can support continued company growth.
Reliability & Observability
- Own and improve the reliability, availability, performance, and scalability of production systems.
- Develop monitoring, alerting, logging, and observability strategies across our infrastructure and applications.
- Lead incident response, troubleshooting, and root‑cause analysis for production issues.
- Establish and improve operational practices around incident management, postmortems, and preventative remediation.
- Identify reliability risks and proactively improve system resilience, capacity, and performance.
- Help define and monitor appropriate SLIs, SLOs, and operational metrics
.
Security & Compliance
- Implement cloud and application security best practices across infrastructure and production systems.
- Partner with Security and Engineering to support HIPAA, SOC 2, and other compliance requirements
. - Implement appropriate controls around access management, secrets, encryption, logging, and infrastructure security.
- Help identify and remediate infrastructure and application security risks.
Technical Leadership
- Own technical initiatives from design through production.
- Partner closely with Software Engineering, Product, Security, and other teams to solve complex technical problems.
- Mentor engineers and contribute to a strong engineering culture.
- Help establish engineering and operational best practices as the company scales.
- Balance reliability, security, engineering velocity, and business priorities when making technical decisions.
- 7+ years of professional software engineering experience
, with significant experience in SRE, platform engineering, Dev Ops, or cloud infrastructure. - Bachelor’s degree…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).