×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer (SRE) – Production Services

Job in Pittsburgh, Allegheny County, Pennsylvania, 15201, USA
Listing for: Apolis
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below

Site Reliability Engineer (SRE) – Production Services

We are seeking an experienced Site Reliability Engineer (SRE) – Production Services to support and enhance production operations through automation, reliability engineering, observability, and self-healing capabilities. The ideal candidate will have strong expertise in Java Spring Boot, Apache Kafka, Dev Ops, and CI/CD automation, along with experience building scalable, resilient, and highly available production systems. This is a fully onsite role in Pittsburgh, PA, and only local candidates will be considered.

Role and Responsibilities:
  • Automate high-volume production support requests and operational workflows.
  • Develop self-service and agent-driven automation solutions to minimize manual effort.
  • Implement standardized operational processes with auditability and resilience.
  • Build auto-retry, backoff, and recovery mechanisms for recurring production failures.
  • Define, monitor, and maintain Service Level Objectives (SLOs) and apply error budget principles.
  • Improve reliability of batch processing through standardized recovery patterns.
  • Develop observability dashboards for incidents, failures, automation coverage, and operational metrics.
  • Create and enhance production runbooks and convert them into automated remediation workflows.
  • Drive permanent resolution of recurring production issues through root cause analysis.
  • Implement self-healing capabilities to reduce operational intervention.
  • Optimize monitoring and alerting platforms (Moogsoft or similar) to improve signal-to-noise ratio.
  • Leverage automation and AI-driven operational solutions for recurring production issues.
  • Collaborate with development, infrastructure, and operations teams to improve system reliability and production stability.
Required Skills:
  • 12+ years of overall IT experience.
  • 8–10+ years of Site Reliability Engineering (SRE) or Production Support experience.
  • Strong hands-on experience with Java and Spring Boot.
  • Experience with Apache Kafka.
  • Strong knowledge of Dev Ops practices and tools.
  • Expertise in CI/CD automation (Jenkins, Git Lab CI, Azure Dev Ops, etc.).
  • Experience with production monitoring, observability, dashboards, and alerting tools.
  • Knowledge of Service Level Objectives (SLOs), SLIs, and Error Budgets.
  • Experience implementing automation, self-healing, and operational runbooks.
  • Strong troubleshooting and root cause analysis skills.
  • Experience working in enterprise production support environments.
Qualifications:
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
  • Experience with cloud platforms and container technologies is a plus.
  • Excellent communication and collaboration skills.
  • Ability to work in a fast-paced production support environment.
  • Local candidates available to work onsite in Pittsburgh, PA.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary