×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer; SRE

Job in Atlanta, Fulton County, Georgia, 30383, USA
Listing for: Convergys
Full Time position
Listed on 2026-09-02
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Salary/Wage Range or Industry Benchmark: 115000 - 140000 USD Yearly USD 115000.00 140000.00 YEAR
Job Description & How to Apply Below
Position: Senior Site Reliability Engineer (SRE)
Job Summary:

We are seeking a Senior Site Reliability Engineer (SRE) to build, automate, and support highly scalable cloud-native platforms and digital applications. This role is responsible for improving system reliability, observability, performance, and operational excellence through Infrastructure as Code, Kubernetes, CI/CD automation, monitoring, and incident management. The ideal candidate will have strong experience with AWS/Azure, distributed systems, platform engineering, and enterprise observability tools, with a passion for reducing operational toil and enhancing platform resilience through automation and intelligent operations.

Responsibilities:

Reliability Engineering & Operational Excellence Design, implement, and support highly available, scalable, and resilient cloud-native platforms and services.

Define, monitor, and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.

Identify, analyze, and eliminate reliability and performance bottlenecks across distributed systems.

Lead incident response, Root Cause Analysis (RCA), and implementation of preventive measures.

Participate in on-call support and production operations for mission-critical applications and services.

Observability & Monitoring Build and manage enterprise observability solutions leveraging tools such as Splunk, Grafana, Prometheus, Datadog, New Relic, and Open Telemetry.

Develop dashboards, alerts, monitoring strategies, and reporting frameworks that provide actionable operational insights.

Improve Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) through proactive monitoring and automation.

Establish best practices for logging, distributed tracing, application monitoring, and system health management.

Platform Engineering Develop and enhance shared platform services and infrastructure utilized by multiple engineering teams.

Create self-service capabilities and automation tools that improve developer productivity and reduce operational overhead.

Design reusable platform components, frameworks, and operational tooling.

Collaborate with architecture, engineering, and product teams to define and execute platform modernization strategies.

Cloud Infrastructure & Automation Design, deploy, and manage cloud infrastructure across AWS and Azure environments.

Implement Infrastructure as Code (IaC) using Terraform, Cloud Formation, Helm, and Kubernetes manifests.

Automate infrastructure provisioning, configuration management, deployments, scaling, and recovery processes.

Improve infrastructure consistency, governance, security, and scalability through automation and standardization.

CI/CD & Dev Ops Enablement Build and optimize CI/CD pipelines to enable secure, scalable, and reliable software delivery.

Implement deployment automation, release management processes, and validation controls.

Support Git Ops methodologies and continuous delivery practices.

Partner with development teams to improve deployment frequency, quality, and operational stability.

Incident Management & Resilience Engineering Lead operational readiness assessments, disaster recovery planning, and business continuity exercises.

Develop and maintain runbooks, playbooks, escalation procedures, and automated remediation workflows.

Drive resiliency testing, chaos engineering initiatives, and fault-tolerance improvements.

Ensure compliance with operational, security, and reliability standards.

AI-Driven Operations & Innovation Utilize AI-powered observability and incident management platforms to enhance operational efficiency.

Leverage predictive analytics and automation to proactively identify risks and prevent service disruptions.

Drive adoption of intelligent operational capabilities that improve system reliability, engineering productivity, and customer experience.

Explore innovative approaches to reduce manual effort and operational toil through automation and AI-driven solutions.

Cross-Functional Collaboration Partner closely with Software Engineering, Platform Engineering, Cloud Architecture, Security, and Product teams to deliver reliable and scalable solutions.

Provide technical leadership and guidance on reliability best practices, automation strategies, and operational excellence.

Mentor junior engineers and contribute to continuous improvement initiatives across the organization.

Qualifications5+ years of experience in Site Reliability Engineering (SRE), Dev Ops, Platform Engineering, Cloud Operations, or a related field.

Strong experience with Linux administration, Kubernetes, Docker, and cloud platforms such as AWS and/or Azure.

Hands-on experience with Infrastructure as Code (Terraform, Cloud Formation, Helm) and CI/CD pipeline automation.

Proficiency in scripting and automation using Python and Bash.

Experience supporting distributed systems, APIs, microservices, and customer-facing applications in production environments.

Strong knowledge of monitoring and observability tools such as Splunk, Grafana, Prometheus, Open…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary