×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineering Manager

Job in San Diego, San Diego County, California, 92189, USA
Listing for: NationsBenefits, LLC
Full Time position
Listed on 2026-08-02
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below

Nations Benefits is recognized as one of the fastest-growing companies in America and a Healthcare Fintech provider of supplemental benefits, flex cards, and member engagement solutions. We partner with managed care organizations to provide innovative healthcare solutions that drive growth, improve outcomes, reduce costs, and bring value to their members.

Through our comprehensive suite of innovative supplemental benefits, fintech payment platforms, and member engagement solutions, we help health plans deliver high-quality benefits to their members that address the social determinants of health and improve member health outcomes and satisfaction.

Our compliance-focused infrastructure, proprietary technology systems, and premier service delivery model allow our health plan partners to deliver high-quality, value-based care to millions of members.

We offer a fulfilling work environment that attracts top talent and encourages all associates to contribute to delivering premier service to internal and external customers alike. Our goal is to transform the healthcare industry for the better! We provide career advancement opportunities from within the organization across multiple locations in the US, South America, and India.
Location: Remote (US-based candidates only)
Manager, Site Reliability Engineering (SRE)
Position Overview

We are seeking a Manager, Site Reliability Engineering (SRE) to lead our US-based SRE team and drive operational excellence across our production platforms.

This is a player-coach leadership role that combines people management with hands‑on technical leadership. You will mentor and grow a team of Site Reliability Engineers while actively participating in major incident response, reliability initiatives, and operational reviews. The role is a key part of our global follow-the-sun support model and requires close collaboration with SRE leadership in India.

Key Responsibilities
Team Leadership & Development
  • Lead, mentor, and develop a US-based team of Site Reliability Engineers.
  • Conduct regular 1:1s, performance reviews, and career development discussions.
  • Own hiring, onboarding, and retention efforts as the team scales.
  • Foster a culture of ownership, blameless postmortems, and continuous improvement.
Operational Excellence & Incident Management
  • Lead day-to-day production operations and ensure timely incident triage, resolution, and escalation.
  • Serve as an escalation point and incident commander for major production incidents.
  • Drive problem management and root cause analysis processes.
  • Carry Pager Duty on-call escalation responsibilities for critical issues.
  • Track and report operational KPIs, SLAs, and SLOs, including availability, MTTR, and incident trends.
Reliability & Automation
  • Improve system reliability, observability, and resilience using Datadog and related tooling.
  • Drive automation, self‑healing capabilities, and runbook maturity.
  • Partner with Development, Dev Ops, Dev Sec Ops , and Engineering teams to embed reliability into the SDLC.
  • Contribute hands‑on to tooling, automation, and technical reviews as needed.
Collaboration & Global Alignment
  • Coordinate closely with SRE leadership in India to ensure seamless follow-the-sun coverage.
  • Represent the US SRE organization in cross‑functional planning and operational reviews.
  • Communicate effectively with both technical and non‑technical stakeholders.
Documentation & Compliance
  • Maintain high‑quality documentation for incidents, postmortems, runbooks, and operational procedures.
  • Ensure adherence to healthcare and fintech compliance standards, including HIPAA, PCI DSS, SOC 2, ISO 27001, and HITRUST.
Required Qualifications
  • 5–8 years of experience in Site Reliability Engineering, Dev Ops, Production Support, or Platform Engineering.
  • 1–2+ years of experience leading, mentoring, or managing engineers.
  • Demonstrated success operating in a player‑coach leadership model.
  • Strong hands‑on experience with production incident management and escalation processes.
  • Proficiency with Datadog or similar observability platforms.
  • Hands‑on experience with Kubernetes and Docker in production environments.
  • Strong scripting or programming skills in Power Shell, Bash, Python, Java,…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary