More jobs:
Senior Site Reliability Engineer, Observability , NY
Job in
New York, New York County, New York, 10261, USA
Listed on 2026-07-19
Listing for:
Ripple
Full Time
position Listed on 2026-07-19
Job specializations:
-
IT/Tech
SRE/Site Reliability
Job Description & How to Apply Below
Location: New York
Senior Site Reliability Engineer, Observability
New York, NY, United States
What you’ll do:This is an engineering-first role with a coaching dimension—not the other way around. You will spend the majority of your time doing hands‑on observability and reliability engineering work: building instrumentation, designing alert configurations, authoring Terraform, and troubleshooting production systems. Alongside that, you will coach and consult with stream-aligned product teams, helping them build operational maturity over time.
Observability Engineering- Design and implement monitoring, alerting, and dashboards in New Relic (APM, Infrastructure, Logs, Synthetics) across Azure and AWS; write NRQL queries for troubleshooting, analysis, and reporting.
- Define and implement SLOs/SLIs and error budgets; coach teams on using them to balance feature velocity with reliability and communicate system health to stakeholders.
- Lead alert noise reduction and signal quality engineering—tune thresholds, eliminate false positives, and ensure every alert is actionable.
- Optimize observability costs through log ingestion management, pipeline rules, and New Relic configuration governance.
- Partner with engineering teams to improve observability maturity: structured logging, metrics instrumentation (RED/USE methods), distributed tracing, and effective dashboard patterns.
- Develop and maintain Terraform infrastructure as code for provisioning and managing monitoring resources, alert configurations, and observability infrastructure—this is a primary engineering responsibility, not an occasional task.
- Establish and enforce IaC governance standards for observability infrastructure across teams, providing a repeatable, auditable model for how monitoring resources are managed.
- Author and troubleshoot Azure Dev Ops pipelines; support teams with deployment visibility, change tracking, and release hygiene as it relates to production reliability.
- Administer and configure Incident.
IO: alert routing, notification workflows, Slack and Ops Genie integration, and runbook management—operationalizing what exists today and expanding from there. - Build out incident management foundations that are largely yours to establish: PIR/postmortem processes, on‑call rotation design, escalation policies, incident severity classification, and response playbooks.
- Track and report on MTTR, MTTD, and incident frequency; identify trends and drive continuous improvement in partnership with engineering teams.
- Respond to and debrief on production incidents—providing real‑time troubleshooting support and facilitating structured post‑incident reviews.
- Enable stream‑aligned engineering teams to adopt improved observability and incident management practices through workshops, consultation, and hands‑on guidance.
- Collaborate with the Subsystems Platform Team to translate common needs into self‑service observability and incident management capabilities.
- Build lasting team competency through documentation, training materials, and knowledge‑sharing sessions that outlast any individual engagement.
- 7+ years in Site Reliability Engineering, Dev Ops, or Platform Engineering with a strong focus on observability and production operations.
- Proven ability to deliver hands‑on engineering work while coaching and mentoring teams—comfortable switching between builder and consultant modes.
- Experience working in Agile/Scrum environments and collaborating effectively with cross‑functional teams.
- Expert‑level hands‑on experience with New Relic (APM, Infrastructure, Logs, Synthetics, Alerts) and strong NRQL proficiency for troubleshooting and analysis.
- Deep understanding of structured logging, metrics collection (RED/USE methods), distributed tracing, and designing effective dashboards and alerts.
- Expertise defining and implementing SLOs/SLIs and error budgets for reliability management.
- Hands‑on experience with incident management platforms (Incident.IO, Pager Duty, Ops Genie, or similar).
- Experience designing incident response workflows, on‑call rotations, escalation policies, and…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×