×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer; Observability NY

Job in New York, New York County, New York, 10261, USA
Listing for: Tech Mirrors
Full Time position
Listed on 2026-08-30
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Support
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below
Position: Site Reliability Engineer (Observability) : New York, NY
Location: New York

Job Description — Site Reliability Engineer (Observability)

Job Location:

New York City, NY (Hybrid — Onsite Required)

Job Type: Long-Term Contract

Client – Gspann / End client not disclosed at this moment.

Key

Skills:

Retail domain with skills like Splunk, Logic Monitor/Dynatrace/Datadog tool capabilities.

Role Summary

We are seeking an experienced Site Reliability Engineer (SRE) with strong Observability expertise to support large-scale, customer-facing retail platforms. The ideal candidate has hands-on experience with enterprise monitoring tool chains and a track record of ensuring system reliability and uptime in high-traffic retail/eCommerce environments. This is a hybrid, long-term contract role based in NYC with required onsite attendance.

Key Responsibilities

Design and maintain observability pipelines covering logs, metrics, traces, and synthetic monitoring across production and non-production environments.

Build dashboards, alerts, and SLO/SLI frameworks using Splunk and one or more of Logic Monitor, Dynatrace, or Datadog.

Partner with application, infrastructure, and platform teams to define monitoring coverage for retail-critical systems (order management, checkout, inventory, POS, fulfillment).

Lead incident response and root cause analysis, using observability data to reduce MTTD and MTTR.

Establish proactive alerting strategies that reduce noise while maintaining signal fidelity for critical services.

Support high-traffic seasonal events (holiday peak, promotions, flash sales) with readiness reviews and real-time monitoring support.

Automate observability configuration (monitoring-as-code) and integrate observability into CI/CD pipelines.

Document runbooks, escalation paths, and post-incident reviews.

Required Skills & Experience

7+ years in SRE, Dev Ops, or Observability/Monitoring engineering roles.

Hands-on expertise with Splunk (search, dashboards, alerting, log pipeline management).

Practical experience with at least one of:
Logic Monitor, Dynatrace, or Datadog.

Prior experience supporting retail or eCommerce platforms, ideally including peak-load event support.

Strong understanding of distributed systems, microservices, and cloud infrastructure (AWS/Azure/GCP).

Experience with incident management, on-call rotations, and postmortem/RCA processes.

Scripting proficiency (Python, Shell, or similar) for monitoring/alerting automation.

Familiarity with containerized environments (Kubernetes, Docker) and CI/CD tooling.

Solid understanding of SLIs, SLOs, error budgets, and reliability engineering principles.

Preferred

Additional observability tools (New Relic, Grafana, Prometheus, ELK).

APM instrumentation experience (Open Telemetry, tracing frameworks).

Infrastructure-as-code experience (Terraform, Ansible).

Familiarity with retail systems: POS, OMS, WMS, inventory/fulfillment platforms.

Relevant cloud or observability tool certifications.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary