×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer III

Job in Irvine, Orange County, California, 92713, USA
Listing for: Socket.dev
Full Time position
Listed on 2026-08-03
Job specializations:
  • IT/Tech
    IT Support, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 117000 - 146500 USD Yearly USD 117000.00 146500.00 YEAR
Job Description & How to Apply Below

Who is Taco Bell?

Taco Bell was born and raised in California and has been around since 1962. We went from selling everyone’s favorite Crunchy Tacos on the West Coast to a global brand with 7,500+ restaurants, 350 franchise organizations, that serve 42+ million fans each week around the globe. We’re not only the largest Mexican-inspired quick service brand (QSR) in the world, we’re also part of the biggest restaurant group in the world:
Yum! Brands. Much of our fan love and authentic connection with our communities are rooted in being rebels with a cause. From ensuring we use high-quality, sustainable ingredients to elevating restaurant technology in ways that haven’t been done before… we will continue to be inclusive, bold, challenge the status quo and push industry boundaries. We’re a company that celebrates and advocates for different, has bold self-expression, strives for a better future, and brings the fun while we’re  fuel our culture with real people who bring unique experiences.

We inspire and enable our teams and the world to Live Más.

About

The Role

At Taco Bell, we’re Cultural Rebels. Want to join in on the passion-fueled fun? Learn more about the career below. We’re looking for someone who can own a problem from start to finish — someone who is comfortable digging into the code, identifying the issue, and proposing a fix. This person will become a subject matter expert within the Taco Bell Digital space.

Fundamentally, we’re looking for someone who lives and breathes observability and automation, and who is passionate about continuously raising the bar. We want someone who listens well, communicates clearly and effectively, drives alignment, and helps engineering teams improve their processes. As a fierce advocate for the customer, your job would take you into troubleshooting issues and incidents, building out new dashboards and alerts, finding out the answers to help us get to the root cause of problems, and ultimately fixing them for good.

The

Day-to-Day
  • Create, update, or automate internal business processes or tools to reduce toil and improve team productivity.
  • Understand and monitor the Taco Bell Digital ecosystem for performance, availability, and accuracy of transactional data.
  • Communicate and collaborate with both technical and non-technical stakeholders on issues, upcoming changes, and updates to system health.
  • Perform final validation tests on various mobile and web-based applications, reporting on, and offering feedback on areas for improvement.
  • Build expertise in serverless infrastructure and initiatives while also learning aspects of modern SRE practices and terms, such as SLIs, SLOs, Observability, toil, and incident response with blameless postmortems.
  • Work with and adopt Agile practices while participating in a 24/7 on-call rotation.
  • Collaborate within the team and with cross-functional partners on high-impact business issues that affect revenue and brand reputation.
Requirements
  • Bachelor’s degree in computer science, engineering, OR a related field, OR equivalent work experience.
  • At least 2+ years of experience in the SRE space, with a focus on observability and automation
  • Familiarity with SRE core principles (e.g. SLO, SLA, SLI, Error Budget, etc.)
  • Hands-on experience creating monitors, dashboards, SLOs, and other observability capabilities
  • Experience with logging solutions or platforms such as Data Dog, Cloud Watch log insights, etc.
  • Understanding of incident management practices, including leading bridge calls, conducting RCAs, and facilitating postmortems
  • Familiarity with modern observability practices and tools, such as distributed tracing, APM, Open Telemetry, and RUM
  • Excellent communication and collaboration skills, with the ability to work effectively in a fast-paced environment as a member of a team
  • A fundamentally complete understanding of Observability principles (not just monitoring) + experience using tools like Data Dog, Lumigo, Cloud Watch, or similar
  • General level understanding of Agile methods such as Kanban, Scrum, etc.
  • Advanced troubleshooting skills
  • A curious mindset and the desire to always keep learning
  • Proactive self-starter capable of operating autonomously
  • Abili…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary