×
Register Here to Apply for Jobs or Post Jobs. X

Senior​/Staff Site Reliability Engineer

Job in New York, New York County, New York, 10261, USA
Listing for: Sage49
Full Time position
Listed on 2026-08-02
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 175000 - 230000 USD Yearly USD 175000.00 230000.00 YEAR
Job Description & How to Apply Below
Location: New York

About this Role

Sage provides life‑saving functionality that improves the lives of our older population. This role is critical to ensure Sage can live up to its mission to be a 24x7, highly available platform for elder care. As a Site Reliability Engineer, you’ll partner with engineering teams across the organization to achieve four 9s of uptime for our platform.

Responsibilities
  • Design and evolve highly reliable system architectures
    , ensuring high availability, fault tolerance, and scalability across Sage’s production infrastructure.
  • Lead complex incident response efforts
    , coordinating across engineering teams to quickly diagnose and resolve production issues while driving thorough post‑incident reviews and long‑term reliability improvements.
  • Define and implement organization‑wide observability practices
    , including metrics, logging, tracing, and actionable alerting to ensure strong visibility into system health.
  • Establish and maintain reliability standards
    , including defining SLIs, SLOs, and error budgets, and partnering with engineering teams to integrate these practices into the software development lifecycle.
  • Drive automation and infrastructure improvements that reduce operational toil and improve the efficiency and reliability of deployments, monitoring, and operational workflows.
  • Partner with engineering teams on system design and architecture reviews
    , ensuring reliability, scalability, and operational best practices are considered early in the development process.
  • Evolve Sage’s cloud infrastructure
    , including networking, compute, storage, and security practices to support scalable and resilient systems.
  • Operate and improve critical data infrastructure
    , ensuring high availability, performance, backup strategies, and disaster recovery processes for production databases.
  • Lead capacity planning and auto‑scaling efforts
    , ensuring infrastructure and systems scale effectively as product usage grows.
  • Build internal tooling and platforms that improve the developer experience, simplify debugging, and enable safer and more reliable deployments.
Qualifications
  • 7-12+ years of experience in software engineering, infrastructure engineering, or site reliability engineering, operating large‑scale distributed systems in production.
  • Experience operating and supporting edge or device‑based systems, including managing connectivity, observability, remote updates, and reliability for distributed hardware deployments such as IoT or field devices.
  • Strong networking fundamentals, including experience debugging distributed system issues across load balancers, DNS, TLS, and VPC networking within platforms like Amazon Virtual Private Cloud or similar cloud networking environments.
  • Experience operating and scaling production databases, including performance tuning, replication, backup/recovery strategies, and high availability for systems such as PostgreSQL, MySQL, or distributed databases.
  • Deep expertise in cloud infrastructure, such as Amazon Web Services or Google Cloud Platform.
  • Strong experience designing and operating highly available systems, including strategies for redundancy, failover, disaster recovery, and capacity planning.
  • Expertise in containerization and orchestration, particularly with Kubernetes and modern container platforms.
  • Advanced observability and monitoring skills, using tools such as Datadog, Prometheus or Grafana.
  • Strong programming ability in languages commonly used for infrastructure and reliability engineering (e.g., Go, Python, or Java), with experience building internal tooling and automation.
  • Deep knowledge of infrastructure‑as‑code practices, including tools like Terraform or Pulumi. Proven experience leading reliability initiatives, such as defining SLOs/SLIs, improving incident response processes, and driving post‑incident reviews.
  • Ability to influence engineering teams across the organization, guiding best practices for reliability, scalability, and operational excellence.
  • Strong incident management and production debugging skills, with experience coordinating responses to complex outages and improving long‑term system resilience.
Preferred Qualifications
  • Experience introducing and scaling SRE…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary