×
Register Here to Apply for Jobs or Post Jobs. X

Senior Software Engineer – SRE & AIOps

Job in Santa Clara, Santa Clara County, California, 95053, USA
Listing for: Servicenow
Full Time position
Listed on 2026-09-12
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 190900 - 334100 USD Yearly USD 190900.00 334100.00 YEAR
Job Description & How to Apply Below

Company Description

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, Service Now is the AI control tower for business reinvention. Our Service Now AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better.

We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

Job Description

About the Role

Service Now is seeking a Senior Staff Reliability Engineer – SRE & AIOpsto drive infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, this technical leader will design and implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.

This role combines deep hands‑on technical expertise in Kubernetes, cloud platforms, and Dev Ops practices with strategic influence across infrastructure teams. You will architect SRE tooling, develop auto‑remediation capabilities, and establish patterns that allow Service Now's cloud platform to maintain industry‑leading reliability while minimizing operational toil across follow‑the‑sun global teams.

What you get to do in this role:

  • Design, deploy, and operate enterprise‑scale Kubernetes clusters across hybrid and multi‑cloud environments, establishing governance, scaling policies, and operational practices that support high‑velocity application deployments %+ availability targets.
  • Architect and implement closed‑loop auto‑remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and machine learning frameworks to predict failures, trigger preventive actions, and continuously reduce MTTR and on‑call burden.
  • Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations, that support global follow‑the‑sun on‑call operations and enable data‑driven incident response.
  • Establish SLO frameworks, error budgets, and alerting policies that balance rapid incident response with alert fatigue management, while developing automated runbooks and playbooks that empower on‑call engineers to resolve issues autonomously.
  • Design and maintain Infrastructure‑as‑Code frameworks and Git Ops pipelines that enable reproducible, auditable infrastructure deployments across hybrid and multi‑cloud environments with consistent security and compliance guardrails.
  • Architect hybrid cloud and data center operations, spanning on‑premises infrastructure, public cloud environments, and edge computing, including workload migration strategies, disaster recovery patterns, and cost optimization practices across multi‑region deployments.
  • Drive adoption of containerization, microservices, and Dev Ops patterns across engineering teams, establishing CI/CD best practices, service mesh architectures, and network security controls that enable rapid, safe release cycles.
  • Design on‑call rotation schedules, escalation policies, and incident command systems that span across different time zones, ensuring 24/7 incident response while driving post‑incident review processes that capture learning and drive systemic improvements.
  • Mentor and guide junior SRE engineers, infrastructure teams, and Dev Ops practitioners…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary