Site Reliability Engineer (SRE) – Production Services
Job in
Pittsburgh, Allegheny County, Pennsylvania, 15201, USA
Listed on 2026-08-05
Listing for:
Apolis
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
SRE/Site Reliability, Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
Site Reliability Engineer (SRE) – Production Services
We are seeking an experienced Site Reliability Engineer (SRE) – Production Services to support and enhance production operations through automation, reliability engineering, observability, and self-healing capabilities. The ideal candidate will have strong expertise in Java Spring Boot, Apache Kafka, Dev Ops, and CI/CD automation, along with experience building scalable, resilient, and highly available production systems. This is a fully onsite role in Pittsburgh, PA, and only local candidates will be considered.
Role and Responsibilities:- Automate high-volume production support requests and operational workflows.
- Develop self-service and agent-driven automation solutions to minimize manual effort.
- Implement standardized operational processes with auditability and resilience.
- Build auto-retry, backoff, and recovery mechanisms for recurring production failures.
- Define, monitor, and maintain Service Level Objectives (SLOs) and apply error budget principles.
- Improve reliability of batch processing through standardized recovery patterns.
- Develop observability dashboards for incidents, failures, automation coverage, and operational metrics.
- Create and enhance production runbooks and convert them into automated remediation workflows.
- Drive permanent resolution of recurring production issues through root cause analysis.
- Implement self-healing capabilities to reduce operational intervention.
- Optimize monitoring and alerting platforms (Moogsoft or similar) to improve signal-to-noise ratio.
- Leverage automation and AI-driven operational solutions for recurring production issues.
- Collaborate with development, infrastructure, and operations teams to improve system reliability and production stability.
- 12+ years of overall IT experience.
- 8–10+ years of Site Reliability Engineering (SRE) or Production Support experience.
- Strong hands-on experience with Java and Spring Boot.
- Experience with Apache Kafka.
- Strong knowledge of Dev Ops practices and tools.
- Expertise in CI/CD automation (Jenkins, Git Lab CI, Azure Dev Ops, etc.).
- Experience with production monitoring, observability, dashboards, and alerting tools.
- Knowledge of Service Level Objectives (SLOs), SLIs, and Error Budgets.
- Experience implementing automation, self-healing, and operational runbooks.
- Strong troubleshooting and root cause analysis skills.
- Experience working in enterprise production support environments.
- Bachelor's degree in Computer Science, Information Technology, Engineering, or related field.
- Experience with cloud platforms and container technologies is a plus.
- Excellent communication and collaboration skills.
- Ability to work in a fast-paced production support environment.
- Local candidates available to work onsite in Pittsburgh, PA.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×