×
Register Here to Apply for Jobs or Post Jobs. X

Senior Platform Reliability Engineer

Job in New York, New York County, New York, 10261, USA
Listing for: Grow Therapy
Full Time position
Listed on 2026-06-26
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 182000 - 250000 USD Yearly USD 182000.00 250000.00 YEAR
Job Description & How to Apply Below
Location: New York

Grow Therapy is on a mission to serve as the trusted partner for therapists growing their practice, and patients accessing high-quality care. Powered by technology, we are a three-sided marketplace that empowers providers, augments insurance payors, and serves patients. Following the mass increase in depression and anxiety, the need for accessibility is more important than ever. To make our vision for mental healthcare a reality, we’re building a team of entrepreneurs and mission-driven go-getters.

Since launching in February 2021, we’ve empowered more than ten thousand therapists and hundreds of thousands of clients across the country and insurance landscape. We’ve raised more than $328

Mm in funding, including our Series D, at a $3B valuation from Sequoia Capital, Transformation Capital, TCV, Signal Fire, Menlo Ventures, Goldman Sachs Alternatives, and others.

About the Role

We’re hiring a Senior Platform Reliability Engineer to help define and scale reliability as a first-class capability  this role you’ll operate horizontally across the organization, shaping how reliability is understood, measured, and built into the developer experience.

You’ll work closely with other members of the platform team as well as our product engineering teams to establish standards around observability, SLOs/SLAs, and incident response—while also helping translate those standards into self-service tooling and “golden paths” that make it easy for teams to adopt them.

This is a high-impact, highly autonomous role where you’ll drive both cultural and technical change, ultimately enabling teams to independently build and operate reliable systems at scale.

What You'll Work On

You’ll help us establish and scale reliability as a discipline at Grow by:

  • Defining Reliability Standards Establishing frameworks for SLOs/SLAs, error budgets, and operational readiness; helping teams understand what to measure and why it matters.

  • Improving Observability & Measurement Identifying gaps in metrics, logging, and tracing; ensuring services are measurable, debuggable, and aligned with reliability goals.

  • Evolving Incident Response Developing and improving incident response practices, from detection to post-incident learning, and helping teams build sustainable on-call and escalation patterns.

  • Enabling Self-Service Reliability Partnering with the platform team to build tooling and abstractions (e.g., service scorecards, dashboards, templates, golden paths) that make it easy for teams to adopt and stay compliant with reliability standards.

  • Driving Adoption Across Teams Working cross-functionally to educate, influence, and guide engineering teams—scaling reliability practices through a combination of clear standards, strong communication, and developer-friendly systems

Who You Are

  • Experienced in production systems: You have 6+ years of experience operating and improving reliability of production systems at scale.

  • Strong foundation in cloud and infrastructure: You have hands‑on experience with AWS, Kubernetes (e.g., EKS), and infrastructure as code tools like Terraform.

  • Deep understanding of reliability principles: You’ve defined or worked with SLOs/SLAs, understand error budgets, and have experience improving reliability through measurement and iteration.

  • Observability expertise: You’ve worked with modern observability tooling (we use Data Dog) and understand how to build actionable monitoring systems across metrics, logs, and traces.

  • Systems thinker: You’re able to zoom out, identify patterns across teams and services, and design solutions that scale beyond a single system.

  • Impact-oriented: You focus on outcomes over output and care deeply about improving real reliability outcomes—not just adding processes.

  • Strong communicator and influencer: You can drive change across teams without direct authority, balancing pragmatism with long-term vision.

  • Self-directed: You thrive in ambiguous environments and are comfortable defining problems, proposing solutions, and executing independently.

  • Team player
    :
    You collaborate well, communicate with empathy, and enjoy mentoring and learning from others.

Bonus Points

  • You’ve helped introduce or scale reliability practices…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary