×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer; SRE

Job in Atlanta, Fulton County, Georgia, 30383, USA
Listing for: Veriipro
Full Time position
Listed on 2026-09-06
Job specializations:
  • IT/Tech
    SRE/Site Reliability, AWS, Cloud Computing: Infrastructure & Operations, Unix/Linux
Salary/Wage Range or Industry Benchmark: 110000 - 160000 USD Yearly USD 110000.00 160000.00 YEAR
Job Description & How to Apply Below
Position: Site Reliability Engineer (SRE)

We are seeking an experienced Site Reliability Engineer (SRE) to support and maintain production systems hosted on AWS. The role focuses on production support, incident management, monitoring, observability, troubleshooting, and improving system reliability and availability.

Core Responsibilities
  • Provide L1/L2 production support for AWS-hosted applications and infrastructure.
  • Monitor and troubleshoot AWS services including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.
  • Triage production incidents, identify root causes, restore services within defined SLAs, and upscale application defects when required.
  • Participate in 24/7 on-call rotations, major incident management, and post-incident reviews.
  • Perform application and infrastructure health checks and proactively address performance, latency, resource utilization, and availability issues.
  • Build and maintain monitoring and observability dashboards using Cloud Watch, Dynatrace, Quantum Metric, and Thousand Eyes.
  • Troubleshoot issues across infrastructure, networking, application, database, and Linux/Unix environments.
  • Support CI/CD pipelines and AWS deployment processes.
  • Apply reliability patterns and continuously improve system stability and operational efficiency.
Required Skills
  • Strong experience supporting AWS production environments.
  • Hands-on experience with incident management and production support.
  • Strong knowledge of Cloud Watch, Dynatrace, Git, observability, and reliability patterns.
  • Experience with monitoring dashboards and operational health checks.
  • Strong troubleshooting and root-cause analysis skills.
  • Working knowledge of CI/CD, databases, and Unix/Linux.
  • Excellent communication and incident coordination skills.
#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary