SRE RunOps Engineer
Listed on 2026-09-04
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Support
7-Eleven is an iconic family of brands with over 86,000 locations, surpassing every retailer in the world.
We revolutionize convenience, restaurants and fuel through cutting edge innovation — working hard to be the customer's first choice. 7-Eleven empowers our employees to 'activate awesome' and make a meaningful impact in their stores and communities every day.
If you're ready to grow, lead and make a difference, come join our team and help shape the future of convenience.
The SRE Run Ops Engineer 2 is responsible for ensuring the reliability, availability, and performance of the 7
NOW delivery platform and associated services through proactive monitoring, incident response, and continuous improvement initiatives. This role combines software engineering principles with operations expertise to maintain high-availability systems that support 7-Eleven's digital delivery business. The SRE Run Ops Engineer 2 collaborates closely with development teams, product managers, and infrastructure engineers to implement automation, optimize system performance, and drive operational excellence. This position plays a critical role in maintaining service level objectives (SLOs) and enhancing the customer experience for 7
NOW users.
DUTIES AND RESPONSIBILITIES:
- Monitor, troubleshoot, and resolve production incidents for 7
NOW platform services, ensuring minimal downtime and rapid restoration of service - Design and implement automation solutions to reduce manual operational tasks and improve system reliability using scripting languages and infrastructure-as-code tools
- Participate in on-call rotation to provide 24/7 support for critical production systems and respond to alerts and incidents
- Collaborate with software engineering teams to improve application observability, implement monitoring solutions, and establish meaningful service level indicators (SLIs)
- Conduct post-incident reviews and root cause analysis to identify systemic issues and implement preventive measures
- Develop and maintain runbooks, operational documentation, and standard operating procedures for common operational scenarios
- Optimize cloud infrastructure costs and resource utilization across AWS or Azure environments supporting the 7
NOW platform - Implement and maintain CI/CD pipelines to enable safe and efficient deployment of application updates
- Perform capacity planning and performance tuning to ensure systems can handle current and projected traffic loads
- Work with security teams to implement and maintain security best practices, vulnerability remediation, and compliance requirements
- Contribute to disaster recovery planning and execute regular testing of backup and recovery procedures
- Mentor junior team members and share knowledge through documentation, training sessions, and code reviews
EDUCATION:
Bachelor's degree in computer science, Information Technology, Engineering, or equivalent practical experience.
YEARS OF RELEVANTWORK EXPERIENCE:
4+
CERTIFICATIONS / LICENSES:AWS or Azure cloud certifications (Solutions Architect Associate, Sys Ops Administrator, or equivalent) preferred
SPECIFIC KNOWLEDGE ANDSKILLS:
- Proficiency in at least one programming or scripting language such as Python, Go, Bash, or Power Shell for automation and tooling development
- Strong understanding of cloud infrastructure services including compute, storage, networking, and managed services in AWS or Azure
- Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, New Relic, or Splunk
- Knowledge of configuration management and infrastructure-as-code tools including Terraform, Ansible, or Cloud Formation
- Solid understanding of Linux/Unix system administration, networking concepts, and distributed systems architecture
- Experience with CI/CD tools and practices including Jenkins, Git Lab CI, Git Hub Actions, or similar platforms.
- Experience utilizing AI to improve Observability, troubleshooting and resolution of production issues.
- Strong troubleshooting and analytical skills with ability to diagnose complex issues across multiple system layers
- Excellent communication skills with ability to clearly document technical processes and explain…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).