Site Reliability Engineer; Observability NY
Listed on 2026-08-30
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Support
Location: New York
Job Description — Site Reliability Engineer (Observability)
Job Location:
New York City, NY (Hybrid — Onsite Required)
Job Type: Long-Term Contract
Client – Gspann / End client not disclosed at this moment.
Key
Skills:
Retail domain with skills like Splunk, Logic Monitor/Dynatrace/Datadog tool capabilities.
We are seeking an experienced Site Reliability Engineer (SRE) with strong Observability expertise to support large-scale, customer-facing retail platforms. The ideal candidate has hands-on experience with enterprise monitoring tool chains and a track record of ensuring system reliability and uptime in high-traffic retail/eCommerce environments. This is a hybrid, long-term contract role based in NYC with required onsite attendance.
Key ResponsibilitiesDesign and maintain observability pipelines covering logs, metrics, traces, and synthetic monitoring across production and non-production environments.
Build dashboards, alerts, and SLO/SLI frameworks using Splunk and one or more of Logic Monitor, Dynatrace, or Datadog.
Partner with application, infrastructure, and platform teams to define monitoring coverage for retail-critical systems (order management, checkout, inventory, POS, fulfillment).
Lead incident response and root cause analysis, using observability data to reduce MTTD and MTTR.
Establish proactive alerting strategies that reduce noise while maintaining signal fidelity for critical services.
Support high-traffic seasonal events (holiday peak, promotions, flash sales) with readiness reviews and real-time monitoring support.
Automate observability configuration (monitoring-as-code) and integrate observability into CI/CD pipelines.
Document runbooks, escalation paths, and post-incident reviews.
Required Skills & Experience7+ years in SRE, Dev Ops, or Observability/Monitoring engineering roles.
Hands-on expertise with Splunk (search, dashboards, alerting, log pipeline management).
Practical experience with at least one of:
Logic Monitor, Dynatrace, or Datadog.
Prior experience supporting retail or eCommerce platforms, ideally including peak-load event support.
Strong understanding of distributed systems, microservices, and cloud infrastructure (AWS/Azure/GCP).
Experience with incident management, on-call rotations, and postmortem/RCA processes.
Scripting proficiency (Python, Shell, or similar) for monitoring/alerting automation.
Familiarity with containerized environments (Kubernetes, Docker) and CI/CD tooling.
Solid understanding of SLIs, SLOs, error budgets, and reliability engineering principles.
PreferredAdditional observability tools (New Relic, Grafana, Prometheus, ELK).
APM instrumentation experience (Open Telemetry, tracing frameworks).
Infrastructure-as-code experience (Terraform, Ansible).
Familiarity with retail systems: POS, OMS, WMS, inventory/fulfillment platforms.
Relevant cloud or observability tool certifications.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).