Site Reliability Engineer
Listed on 2026-08-01
-
IT/Tech
IT Support, Systems Analyst, Cloud Computing: Infrastructure & Operations
Kroger Next Generation Point of Sale (NGPOS) Platform Support Engineer
Level 3 (LV3) Support, Site Reliability, and Operations Engineer supporting Kroger's Next Generation Point of Sale (NGPOS) platform. This team is embedded within the engineering organization, partnering directly with developers throughout the SDLC rather than sitting in a separate support silo.
In office, Blue Ash, OH — 5 days a week; willingness to travel and provide onsite support for new go-lives, pilots, and rollouts. Participate in a scheduled on-call rotation (nights/weekends/holidays) for 24x7 support; PTO/leave may be restricted during peak retail periods and major go-lives.
Bachelor's degree in Computer Science, Information Systems, or a related field (or equivalent experience) preferred; prior experience in application support, operations, or engineering is a plus but not required. Exposure to Point of Sale systems in enterprise environments is helpful; some experience with Java is preferred, Go is a plus.
Any exposure to SQL or No
SQL databases (e.g., MongoDB), scripting (Shell, Power Shell, or Python), or ITSM/monitoring tooling (Service Now, Jira, Confluence, Splunk, Grafana, or equivalent) is a plus — willingness to learn these on the job matters more than prior mastery. Interest in or willingness to learn payment card security standards (PCI-DSS).
Excellent oral and written communication skills; able to translate between business users and Kroger Technology teams. Familiarity with the Agile process is helpful; experience as a Scrum Master a plus. Willingness to speak up and challenge developers on Agile best practices and definitions of done during refinement and QA.
Accountable for driving support tickets and emails to resolution. Good judgment under pressure; able to prioritize across competing support, testing, and scrum-master duties.
Partner with cross-functional teams to expedite issue resolution and manage to SLAs. Support and maintain infrastructure and applications, including off-hours support (24 x
7) as required. Test infrastructure and application changes. Establish priorities and develop functional and programming specifications for application enhancements and modifications.
Participate in the application technical design process; design, code, and unit test application changes using SDLC best practices. Complete estimates and work plans for design, development, implementation, and rollout tasks. Monitor systems, consoles, and performance for service interruptions or delays. Address and resolve infrastructure system failures. Execute, log, and report break/fix changes, service requests, and support activities.
Maintain operational procedures, processes, and scripts; follow documented processes to ensure infrastructure stability. Own incident and problem management: drive major incident response, root cause analysis, and post-incident reviews. Champion engineering standards and continuously improve software delivery processes. Build partnerships across application, business, and infrastructure teams. Independently execute large projects and lead other analysts in completing projects.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).