×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer II; NGPOS Operations Support

Job in Cincinnati, Hamilton County, Ohio, 45208, USA
Listing for: Hudson Manpower
Full Time position
Listed on 2026-07-13
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 105000 - 145000 USD Yearly USD 105000.00 145000.00 YEAR
Job Description & How to Apply Below
Position: Site Reliability Engineer II (NGPOS Operations Support)

Position Overview

We are seeking a hands-on Site Reliability Engineer II to support a next-generation Point of Sale (NGPOS) platform in a highly visible production environment. Unlike traditional SRE roles focused primarily on automation or platform engineering, this position emphasizes production reliability, incident leadership, operational excellence, and engineering support
.

The ideal candidate will lead major incident response efforts, drive Root Cause Analysis (RCA), improve system observability, and collaborate closely with Software Engineering, Platform Engineering, Infrastructure, and Business Operations teams to enhance overall platform reliability.

This role is ideal for someone who enjoys solving complex production issues under pressure while contributing to long-term engineering improvements.

Location: Blue Ash, OH (Cincinnati) – Onsite (5 Days/Week)
Employment Type: W2 – Contract-to-Hire
Duration: Full-Time
Work Authorization: Permanent Residents only. Must be able to convert to full-time without sponsorship.
Experience

Required:

3+ Years

Nice to Have
  • Enterprise Point of Sale (POS) systems

  • Retail technology experience

  • Automation scripting

  • Monitoring optimization

  • Runbook creation

  • Store technology deployments

Key Responsibilities
  • Lead major incident response during production outages

  • Serve as Incident Commander during P1/P2 incidents

  • Coordinate technical bridge calls

  • Communicate outage status to engineering teams and business leadership

  • Lead Root Cause Analysis (RCA) activities

  • Track corrective actions through completion

  • Improve production reliability and system stability

  • Enhance monitoring and observability

  • Reduce alert fatigue

  • Partner with Software Engineering and Platform Engineering teams

  • Support retail store deployments

  • Develop operational documentation, runbooks, and playbooks

  • Participate in after-hours support rotations and maintenance windows

  • Improve service health using SLIs and SLOs

Technical Environment

Monitoring & Observability

  • Dynatrace

  • Azure Monitor

  • Log Analytics

  • Metrics

  • Dashboards

Cloud

  • Microsoft Azure

  • Google Cloud Platform (GCP)

Containers

  • Kubernetes

  • Docker

Operating Systems

  • Linux

Scripting Languages

  • Bash

  • Python

Agile Tools

  • Jira

Enterprise Environment

  • Retail systems

  • Point of Sale (POS)

  • Production Support

  • Hybrid Infrastructure

Ideal Candidate Profile

The ideal candidate will demonstrate:

  • Strong leadership during production incidents

  • Excellent troubleshooting and analytical skills

  • Effective communication under pressure

  • Ownership and accountability

  • Experience coordinating multiple engineering teams

  • Strong operational discipline

  • Continuous improvement mindset

  • Passion for reliability engineering

  • Excellent documentation skills

Required Skills

Must Have

  • Major Incident Management experience

  • Experience serving as Incident Commander

  • Leading production bridge calls

  • Coordinating cross-functional technical teams

  • Executive communication during P1/P2 outages

  • Root Cause Analysis (RCA)

    • Five Whys

    • Fishbone Analysis

    • Timeline reconstruction

    • Corrective action tracking

  • Production troubleshooting across:

    • Cloud environments

    • On‑premise infrastructure

    • Retail/POS systems

  • Observability and Monitoring

    • Dynatrace

    • Azure Monitor

    • Log analysis

    • Dashboards

    • Metrics

  • Linux administration

  • Bash and/or Python scripting

  • Kubernetes

  • Docker

  • Microsoft Azure and/or Google Cloud Platform (GCP)

  • Agile methodology

  • Jira

  • Strong communication and collaboration skills

  • Ability to work onsite five days per week

  • Willingness to travel and participate in on‑call rotations

Travel Requirements
  • Local travel initially throughout Cincinnati and Louisville

  • Travel will expand as additional retail locations are deployed

  • Approximately every four weeks as new sites go live

  • Shared on‑call and travel rotation with a growing six‑person team

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary