Senior Software Engineer – SRE & AIOps
Listed on 2026-09-12
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Company Description
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, Service Now is the AI control tower for business reinvention. Our Service Now AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better.
We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.
Join us to put AI to work for people.
Job DescriptionAbout the Role
Service Now is seeking a Senior Staff Reliability Engineer – SRE & AIOpsto drive infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, this technical leader will design and implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.
This role combines deep hands‑on technical expertise in Kubernetes, cloud platforms, and Dev Ops practices with strategic influence across infrastructure teams. You will architect SRE tooling, develop auto‑remediation capabilities, and establish patterns that allow Service Now's cloud platform to maintain industry‑leading reliability while minimizing operational toil across follow‑the‑sun global teams.
What you get to do in this role:
- Design, deploy, and operate enterprise‑scale Kubernetes clusters across hybrid and multi‑cloud environments, establishing governance, scaling policies, and operational practices that support high‑velocity application deployments %+ availability targets.
- Architect and implement closed‑loop auto‑remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and machine learning frameworks to predict failures, trigger preventive actions, and continuously reduce MTTR and on‑call burden.
- Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations, that support global follow‑the‑sun on‑call operations and enable data‑driven incident response.
- Establish SLO frameworks, error budgets, and alerting policies that balance rapid incident response with alert fatigue management, while developing automated runbooks and playbooks that empower on‑call engineers to resolve issues autonomously.
- Design and maintain Infrastructure‑as‑Code frameworks and Git Ops pipelines that enable reproducible, auditable infrastructure deployments across hybrid and multi‑cloud environments with consistent security and compliance guardrails.
- Architect hybrid cloud and data center operations, spanning on‑premises infrastructure, public cloud environments, and edge computing, including workload migration strategies, disaster recovery patterns, and cost optimization practices across multi‑region deployments.
- Drive adoption of containerization, microservices, and Dev Ops patterns across engineering teams, establishing CI/CD best practices, service mesh architectures, and network security controls that enable rapid, safe release cycles.
- Design on‑call rotation schedules, escalation policies, and incident command systems that span across different time zones, ensuring 24/7 incident response while driving post‑incident review processes that capture learning and drive systemic improvements.
- Mentor and guide junior SRE engineers, infrastructure teams, and Dev Ops practitioners…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).