×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer, Observability , NY

Job in New York, New York County, New York, 10261, USA
Listing for: Ripple
Full Time position
Listed on 2026-07-19
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 160000 - 200000 USD Yearly USD 160000.00 200000.00 YEAR
Job Description & How to Apply Below
Position: Senior Site Reliability Engineer, Observability New York, NY, United States
Location: New York

Senior Site Reliability Engineer, Observability

New York, NY, United States

What you’ll do:

This is an engineering-first role with a coaching dimension—not the other way around. You will spend the majority of your time doing hands‑on observability and reliability engineering work: building instrumentation, designing alert configurations, authoring Terraform, and troubleshooting production systems. Alongside that, you will coach and consult with stream-aligned product teams, helping them build operational maturity over time.

Observability Engineering
  • Design and implement monitoring, alerting, and dashboards in New Relic (APM, Infrastructure, Logs, Synthetics) across Azure and AWS; write NRQL queries for troubleshooting, analysis, and reporting.
  • Define and implement SLOs/SLIs and error budgets; coach teams on using them to balance feature velocity with reliability and communicate system health to stakeholders.
  • Lead alert noise reduction and signal quality engineering—tune thresholds, eliminate false positives, and ensure every alert is actionable.
  • Optimize observability costs through log ingestion management, pipeline rules, and New Relic configuration governance.
  • Partner with engineering teams to improve observability maturity: structured logging, metrics instrumentation (RED/USE methods), distributed tracing, and effective dashboard patterns.
Infrastructure & IaC
  • Develop and maintain Terraform infrastructure as code for provisioning and managing monitoring resources, alert configurations, and observability infrastructure—this is a primary engineering responsibility, not an occasional task.
  • Establish and enforce IaC governance standards for observability infrastructure across teams, providing a repeatable, auditable model for how monitoring resources are managed.
  • Author and troubleshoot Azure Dev Ops pipelines; support teams with deployment visibility, change tracking, and release hygiene as it relates to production reliability.
  • Administer and configure Incident.

    IO: alert routing, notification workflows, Slack and Ops Genie integration, and runbook management—operationalizing what exists today and expanding from there.
  • Build out incident management foundations that are largely yours to establish: PIR/postmortem processes, on‑call rotation design, escalation policies, incident severity classification, and response playbooks.
  • Track and report on MTTR, MTTD, and incident frequency; identify trends and drive continuous improvement in partnership with engineering teams.
  • Respond to and debrief on production incidents—providing real‑time troubleshooting support and facilitating structured post‑incident reviews.
Cross‑Functional Enablement
  • Enable stream‑aligned engineering teams to adopt improved observability and incident management practices through workshops, consultation, and hands‑on guidance.
  • Collaborate with the Subsystems Platform Team to translate common needs into self‑service observability and incident management capabilities.
  • Build lasting team competency through documentation, training materials, and knowledge‑sharing sessions that outlast any individual engagement.
What you’ll bring
  • 7+ years in Site Reliability Engineering, Dev Ops, or Platform Engineering with a strong focus on observability and production operations.
  • Proven ability to deliver hands‑on engineering work while coaching and mentoring teams—comfortable switching between builder and consultant modes.
  • Experience working in Agile/Scrum environments and collaborating effectively with cross‑functional teams.
Observability & Incident Management Expertise — Required
  • Expert‑level hands‑on experience with New Relic (APM, Infrastructure, Logs, Synthetics, Alerts) and strong NRQL proficiency for troubleshooting and analysis.
  • Deep understanding of structured logging, metrics collection (RED/USE methods), distributed tracing, and designing effective dashboards and alerts.
  • Expertise defining and implementing SLOs/SLIs and error budgets for reliability management.
  • Hands‑on experience with incident management platforms (Incident.IO, Pager Duty, Ops Genie, or similar).
  • Experience designing incident response workflows, on‑call rotations, escalation policies, and…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary