Senior Site Reliability Engineer, AI Agents & Automation
Job in
Northern, Floyd County, Kentucky, USA
Listed on 2026-09-16
Listing for:
ServiceTitan, Inc.
Full Time
position Listed on 2026-09-16
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
US Remote:
Full time:
Posted Today:
JR115077
**** Ready to be a Titan?
**** We're looking for a Senior Site Reliability Engineer to join our Site Reliability & Infrastructure Engineering team. We run entirely on the cloud, and this team owns the reliability and health of the applications running on top of it — designing the signals that tell us when something's wrong, and building the systems that keep Service Titan running better, faster, and cheaper as we scale.
We make a huge impact on thousands of companies in the U.S. and abroad by enabling them to be more efficient and effective at running their business. Our Site Reliability and Infrastructure Engineering team centralizes the concerns of measurement and guidance so every engineer can improve availability and efficiency in their own area of the Service Titan cloud. We have a cultural foundation built on diversity, inclusion, and innovation, and we want you and your ideas to thrive e join us.
**** What You'll Do
***** Participate in an on-call rotation, using runbooks and playbooks to diagnose and resolve production issues (e.g., adjusting Horizontal Pod Autoscaler rules in response to load).
* Design, build, and maintain observability dashboards and alerting grounded in Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
* Operate and improve our Kubernetes-based compute platform, which runs the large majority of our infrastructure.
* Work across cloud networking and infrastructure (Azure/AWS) to support reliable, scalable systems.
* Investigate and resolve production incidents, including root-cause analysis and follow-up remediation work.
* Build and operate AI agents and automation that take on manual, repetitive SRE work directly — not just tools that assist a human doing the work.
* Partner with product engineering teams to review architecture and infrastructure decisions before they ship.
* Write and maintain runbooks and documentation so on-call knowledge is shared across the team, not siloed with one person.
* Help define non-functional requirements — scalability, availability, performance — for new systems as they're designed.
* Collaborate across engineering teams to adopt best practices in reliability and observability.
* Contribute to CI/CD pipelines and help teams ship changes safely and quickly.
**** What You'll Bring
***** Kubernetes (must-have): strong, hands-on understanding of Kubernetes as a system.
* AI-native SRE practice (must-have): hands-on, personal experience using AI tools and agents (e.g., Claude Code, Git Hub Copilot Workspace, Cursor, or custom agents built on Claude/MCP servers) to diagnose, automate, and resolve infrastructure and reliability work. +
*** Agentic depth:
*** You've built or operated agents that autonomously monitor infrastructure and take action (e.g., an agent that watches system load and scales, remediates, or escalates without a human in the loop) — not just general AI coding assistance. You can speak concretely to how these agents are actually built and operated (MCP protocol, what a harness is, how to differentiate agent-design strategies rather than just naming tools), to context management (e.g., progressive-disclosure strategies for surfacing the right information without dumping everything into context), and to securing agent actions (scoping authorization, guardrails, what's available in the AI infra ecosystem to enforce it).
+
*** SRE application:
*** That agentic work is pointed at infrastructure and reliability problems specifically — you use AI to move at a materially faster pace in the SRE domain, not as a general-purpose coding aid.
* SRE principles: practical experience with SLIs, SLOs, and error…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×