×
Register Here to Apply for Jobs or Post Jobs. X

Senior Platform Engineering Manager - AI-Driven Reliability

Job in San Francisco, San Francisco County, California, 94199, USA
Listing for: Hinge Health
Full Time position
Listed on 2026-09-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 212000 - 318000 USD Yearly USD 212000.00 318000.00 YEAR
Job Description & How to Apply Below

The Opportunity

Every product surface at Hinge Health—the ones helping 18M+ people move beyond pain—depends on the foundational infrastructure your team will own. As Senior Engineering Manager, Service Enablement, you'll lead the platform-minded team responsible for keeping our cloud infrastructure rock-solid while accelerating the shift to AI-native engineering workflows. Your work will be a force multiplier for every engineer, product manager, and designer in the company.

This is a high-impact leadership role for someone who thrives at the intersection of platform reliability, developer enablement, and autonomous AI-driven operations

What You'll Accomplish

In your first 3 months:

  • Establish relationships with key stakeholders across Security, Product Engineering, Data Engineering, and the India-based SRE team. Audit the current infrastructure posture, on-call practices, and CI/CD pipelines to identify the highest-leverage improvement opportunities.
  • Assess your team of 8–10 engineers (Staff, Senior, and mid-level SREs), understand individual growth trajectories, and begin shaping a coaching-first operating cadence.
In your first 6 months:

  • Own Platform Reliability & Scalability — Guarantee stability and performance across multiple EKS clusters, maintaining 99.9%+ uptime while optimizing cost-efficiency as AI workloads scale.
  • Drive the Harness Engineering Initiative — Define safety rails, test harnesses, and verification systems that allow autonomous AI agents to reliably build, test, and maintain infrastructure in production-grade environments.
  • Scale the NestJS Monorepo & CI/CD Platform — Accelerate migration of all NestJS services into the monorepo, ensuring NX, Git Hub Actions, and Okteto deliver a world-class developer experience.
In your first year:

  • Execute Vendor Strategy & Cost Optimization — Own relationships with AWS, Datadog, Cloudflare, Temporal, and Infisical; deliver significant cost savings as Hinge Health scales infrastructure and AI workloads.
  • Champion Operational Excellence — Run a follow-the-sun support model, maintain 24/7 on-call coverage, lead the Incident Review Board, and continuously improve SLOs, runbooks, and observability to reduce MTTR.
  • Build a High-Performing, Inclusive Team — Develop talent toward Senior and Staff levels, maintain sustainable on-call practices, and foster a culture of blameless retros and continuous improvement.
Who You Are
  • A Platform Thinker — You see infrastructure as a product. You obsess over developer experience metrics (DXI, DORA, PR throughput) and build systems that reduce friction for every team you serve.
  • AI-Forward — You're energized—not threatened—by the shift to AI-assisted engineering. You've experimented with tools like Cursor, Claude Code, or Copilot and can envision a world where engineers act as architects and auditors of AI-generated code.
  • A Trust Builder — You communicate effectively across engineering, security, product, and executive stakeholders to align infrastructure decisions with business priorities.
  • A Learn-it-All — You don't have all the answers, but you have a proven framework for finding them. You lead blameless retros that convert insights into action and measure what matters most.
  • Hands-On & Accountable — You know the ins and outs of your team's tech stack and spend ~15% of your time working in the code. You take owner-operator pride in supporting production systems.
  • Resilient — You thrive in fast-paced, ambiguous environments and can make results happen even with imperfect data—whether that's navigating a breaking-change migration at scale or an unexpected incident.
Basic Requirements
  • At minimum, 3-4 years of professional experience in technology, with depth in SaaS platforms and large-scale distributed systems.
  • At minimum, 3-4 years managing engineering teams of 6+ direct reports
  • Bachelor’s Degree (or equivalent) in Computer Science, Engineering, or a related technical field.
  • Deep hands-on expertise with cloud infrastructure (AWS), container orchestration (Kubernetes/EKS), and Infrastructure as Code (Terraform).
  • Proven track record owning platform reliability at scale — including incident management, SLO-driven operations, and cost…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary