×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer

Job in Austin, Travis County, Texas, 78716, USA
Listing for: Programmers.io
Full Time position
Listed on 2026-07-23
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, IT Support
Salary/Wage Range or Industry Benchmark: 120000 - 180000 USD Yearly USD 120000.00 180000.00 YEAR
Job Description & How to Apply Below

Location:
Austin , TX - Hybrid (4 Days Onsite + 1 Day Remote)

Role

Summary:

Experienced SRE Engineer with 5+ years in designing, managing and supporting distributed systems across multi-cloud environments.

Key Skills & Expertise:

  • Monitoring & Observability:
    Splunk, Grafana, App Dynamics, Thousand Eyes
  • Messaging & Streaming:
    Kafka, MQ
  • Protocols & Web Services: HTTP, DNS, TCP/UDP, REST, SOAP, JSON

Core Competencies:

  • Strong troubleshooting and debugging in microservices architecture
  • Incident management, issue resolution and RCA creation
  • Enterprise cloud infrastructure handling
  • Agile development practices with tools like Git, Jira, Confluence

Responsibilities:

  • Develop and maintain tooling used for environment monitoring and task automation
  • Identify application reliability and availability improvements and build solutions to drive an improved experience
  • Analyze and establish efficient configurations for software and servers, DB connections, indexes, drivers, etc.
  • Coordinate with development teams, technical and non-technical Partners and clients to maintain wide knowledge on dependencies of the critical business transaction including platform, services and tools
  • Monitor internal and vendor service level objectives (SLOs) and agreements (SLAs); identifies and resolves SLO / SLA gaps
  • Serve as technical subject matter expert (SME) for cross-functional engineering Teams;
  • Assist with and troubleshoot systems-related issues and maintenance
  • Collaborate on maintaining services once they are live; measures and monitors availability, latency, and overall system health
  • Develop run book and build automation
  • Develop and maintain E2E monitoring dashboards to support critical business transaction
  • Develop and maintain synthetic monitoring for critical business transaction using tools such as Thousand Eyes
  • Practice sustainable incident response and blameless postmortems
  • Document and promote SRE standards and procedures
  • Develop and assist in deployment and rollback automation
  • Review Release and deployments requirements
  • Build and setup automation tests.
  • Incident communication to impacted stakeholders
  • Coach and mentor junior engineers and fellow practitioners
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary