×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Operations III

Job in Bentonville, Benton County, Arkansas, 72712, USA
Listing for: Walmart
Full Time position
Listed on 2026-09-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability
Salary/Wage Range or Industry Benchmark: 140000 - 180000 USD Yearly USD 140000.00 180000.00 YEAR
Job Description & How to Apply Below
Position: (USA) Site Reliability Operations III

Position Summary…

Role Summary

Walmart is seeking a Site Reliability Operations III team member to assist in driving the reliability, availability, and performance of mission-critical applications and platforms within the Supply Chain & Transportation support area. This role focuses on incident management, monitoring and alerting, troubleshooting complex outages, and applying Dev Ops best practices to maintain stable, scalable systems. The ideal candidate partners closely with engineering and business stakeholders, drives continuous improvement, and upholds Walmart’s values and operational excellence standards.

About the Team

The Walmart Global Tech organization builds and operates the platforms that power Walmart’s stores, digital channels, and supply chain at global scale. Within the Site Reliability Engineering and operations-focused teams, engineers are responsible for proactive system health, rapid incident resolution, and evolving observability and resiliency practices to support a highly available, customer-first ecosystem.

What you’ll do…
  • Incident Management & Reliability
    • Lead and contribute to incident triage, diagnosis, and restoration within defined SLAs.
    • Coordinate with internal teams and external partners to resolve escalated and complex outages.
    • Document incident resolution, troubleshooting steps, and contribute to knowledge management and RCCA activities.
  • Monitoring, Alerting & Observability
    • Define and monitor SLIs, SLOs, and KPIs such as availability, latency, MTBF, and MTTR.
    • Design alerting strategies, thresholds, and instrumentation to improve system awareness and reduce noise. e.g. yaml, python, git.
    • Identify observability gaps and recommend enhancements to monitoring and alerting logic.
  • Triaging & Troubleshooting
    • Independently troubleshoot application performance and availability issues.
    • Perform root cause analysis and contribute data and insights to RCA initiatives.
    • Analyze historical defects to prevent recurrence.
  • Dev Ops & Operational Excellence
    • Execute complex application maintenance, corrective, adaptive, and reengineering activities.
    • Analyze logs
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary