×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer (SRE

Job in Bloomfield, Essex County, New Jersey, 07003, USA
Listing for: Apolis
Full Time position
Listed on 2026-08-02
Job specializations:
  • IT/Tech
    IT Support, Cybersecurity
Job Description & How to Apply Below
Position: Senior Site Reliability Engineer (SRE)
Role Name Senior Site Reliability Engineer (SRE)

Location:

Atlanta US

JOB DESCRIPTION

Role Summary

The Senior Site Reliability Engineer (SRE) is a hands-on role responsible for the availability, performance, and end-to-end observability of QSR digital platforms across Mobile (iOS/Android), Web, and POS systems.

This role is part of the Observability team and works closely with mobile, web, and backend engineering teams to ensure full visibility into customer journeys and user experience.

The focus is on building and operating Real User Monitoring (RUM), synthetic monitoring, and end-to-end telemetry correlation using Splunk and Signal Fx, ensuring issues are detected before customer impact. This is not a monitoring-only role it requires active involvement in instrumentation, release observability, and reliability engineering.

##

Key Responsibilities

- Define and enforce SLIs, SLOs, and error budgets for critical customer journeys (ordering, checkout, payments)
- Own end-to-end observability across Mobile, Web, and POS platforms
- Implement and operate RUM and synthetic monitoring for customer-facing journeys
- Build mobile-first monitoring coverage including app performance, crash rates, API performance, and user journey tracking
- Use Splunk and Signal Fx to design dashboards, detectors, and actionable alerts
- Enable correlation across mobile ? CDN ? API ? backend systems using logs, metrics, and traces
- Partner with engineering teams for instrumentation, SDK integration, and embedding observability into releases
- Analyze telemetry to detect post-release issues, device/OS-specific failures, and network degradation
- Lead response for P1/P2 incidents and drive root cause analysis
- Automate operational toil and improve reliability

## Required Skills

- Strong experience with Splunk (logs, dashboards) and Signal Fx (metrics/APM)
- Hands-on with RUM and synthetic monitoring tools

- Experience with mobile observability (iOS/Android), including performance monitoring and crash analysis
- Strong understanding of distributed systems and microservices (Java, Node.js)

- Experience with Azure (AKS, App Services, APIM)
- Ability to correlate frontend issues with backend services

- Experience with CI/CD pipelines and observability in release processes

## What Success Looks Like

- Clear visibility into mobile, web, and POS customer journeys
- Issues identified before customer complaints or app store feedback
- Strong, low-noise user-impact-driven alerting
- Reduced crash rates, latency, and checkout failures
- Observability embedded into every release
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary