×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer, AI & Agentic Systems

Job in Plano, Collin County, Texas, 75086, USA
Listing for: ServiceLink IP Holding Company, LLC
Full Time, Part Time position
Listed on 2026-07-01
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations
Salary/Wage Range or Industry Benchmark: 100000 - 130000 USD Yearly USD 100000.00 130000.00 YEAR
Job Description & How to Apply Below

Overview

As our SRE charter continues to evolve, this role demands strong hands-on ownership of production reliability and troubleshooting, coupled with advanced capabilities in AI- and agentic-driven automation and performance engineering.

The Site Reliability Engineer will play a critical role in ensuring reliability, scalability, performance, and operational excellence of our platforms. The ideal candidate will leverage Azure-native AI services and agentic systems to reduce toil, improve incident response, and enable intelligent operations—while also driving performance testing practices to validate system resilience under load.

This is a hybrid role, located at our Plano, TX office. Candidates must be willing and able to work in-office 3 days per week in Plano, TX.

Applicants must be currently authorized to work in the United States on a full-time basis and must not require sponsorship for employment visa status now or in the future

A DAY IN THE LIFE

In this role, you will…

  • Own end-to-end reliability of large-scale, Azure-hosted production systems, ensuring high availability, fault tolerance, and graceful degradation
  • Lead hands-on incident troubleshooting, root cause analysis (RCA), and post-incident reviews with actionable follow-ups
  • Build and operate resilient, scalable services on Microsoft Azure (AKS, App Services, Functions, Event Hubs, etc.)
  • Design and maintain comprehensive observability platforms using Prometheus for metrics, Loki for log aggregation, Tempo for distributed tracing, and Grafana for dashboarding and alerting
  • Design, develop, and execute performance testing strategies for distributed systems and microservices, including load testing, stress testing, soak testing, and capacity planning
  • Integrate AI agents with Azure monitoring stack, CI/CD tooling, and incident management platforms
  • Contribute to evolving SRE standards, tooling, operational processes, and knowledge base
Responsibilities

Reliability Engineering & Production Ownership

  • Own end-to-end reliability of large-scale, Azure-hosted production systems, ensuring high availability, fault tolerance, and graceful degradation
  • Lead hands-on incident troubleshooting, root cause analysis (RCA), and post-incident reviews with actionable follow-ups
  • Define, measure, and enforce Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets aligned with business outcomes
  • Drive proactive reliability improvements based on operational insights, failure mode analysis, and capacity planning
  • Participate in on-call rotations and take real-time ownership during production incidents

Platform & Automation Engineering

  • Build and operate resilient, scalable services on Microsoft Azure (AKS, App Services, Functions, Event Hubs, etc.)
  • Design and maintain comprehensive observability platforms using Prometheus for metrics, Loki for log aggregation, Tempo for distributed tracing, and Grafana for dashboarding and alerting
  • Create automation to eliminate manual operational tasks and reduce Mean Time to Recovery (MTTR)
  • Implement self-healing mechanisms, automated remediation workflows, and runbook automation
  • Manage and optimize API lifecycle and traffic management using Gravitee API Gateway
  • Design and implement durable, fault-tolerant workflows and microservice orchestration patterns using Temporal
  • Administer and tune PostgreSQL databases for reliability, performance, and high availability
  • Partner with application and platform teams to improve service operability, deployment safety, and change management

Performance Testing & Load Engineering

  • Design, develop, and execute performance testing strategies for distributed systems and microservices, including load testing, stress testing, soak testing, and capacity planning
  • Build and maintain performance test scripts and virtual user scenarios using Micro Focus Load Runner and VuGen (Virtual User Generator)
  • Analyze performance test results to identify bottlenecks, regressions, and scalability limits; produce clear reports with actionable recommendations
  • Integrate performance testing into CI/CD pipelines to enable continuous performance validation and shift-left testing practices
  • Establish and monitor performance…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary