Site Reliability Operations III
Listed on 2026-09-05
-
IT/Tech
SRE/Site Reliability
Position Summary…
Role Summary
Walmart is seeking a Site Reliability Operations III team member to assist in driving the reliability, availability, and performance of mission-critical applications and platforms within the Supply Chain & Transportation support area. This role focuses on incident management, monitoring and alerting, troubleshooting complex outages, and applying Dev Ops best practices to maintain stable, scalable systems. The ideal candidate partners closely with engineering and business stakeholders, drives continuous improvement, and upholds Walmart’s values and operational excellence standards.
About the Team
The Walmart Global Tech organization builds and operates the platforms that power Walmart’s stores, digital channels, and supply chain at global scale. Within the Site Reliability Engineering and operations-focused teams, engineers are responsible for proactive system health, rapid incident resolution, and evolving observability and resiliency practices to support a highly available, customer-first ecosystem.
What you’ll do…- Incident Management & Reliability
- Lead and contribute to incident triage, diagnosis, and restoration within defined SLAs.
- Coordinate with internal teams and external partners to resolve escalated and complex outages.
- Document incident resolution, troubleshooting steps, and contribute to knowledge management and RCCA activities.
- Monitoring, Alerting & Observability
- Define and monitor SLIs, SLOs, and KPIs such as availability, latency, MTBF, and MTTR.
- Design alerting strategies, thresholds, and instrumentation to improve system awareness and reduce noise. e.g. yaml, python, git.
- Identify observability gaps and recommend enhancements to monitoring and alerting logic.
- Triaging & Troubleshooting
- Independently troubleshoot application performance and availability issues.
- Perform root cause analysis and contribute data and insights to RCA initiatives.
- Analyze historical defects to prevent recurrence.
- Dev Ops & Operational Excellence
- Execute complex application maintenance, corrective, adaptive, and reengineering activities.
- Analyze logs
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).