×
Register Here to Apply for Jobs or Post Jobs. X

Principle SRE Engineer

Job in Atlanta, Fulton County, Georgia, 30301, USA
Listing for: E-Solutions
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below

Job Title

Key Responsibilities

  • Perform SRE operations for distributed systems, ensuring high availability, reliability, and operational excellence.
  • AI in SRE
  • Partner with application/domain teams to strengthen their SRE maturity and operational readiness.
  • Write automation, scripts, and REST APIs to integrate with external systems and eliminate repetitive tasks.
  • Onboard services to Dynatrace/observability platforms; define dashboards, alerts, SLIs, SLOs.
  • Architect and implement resiliency patterns including failover strategies, circuit breakers, graceful degradation.
  • Drive cost optimization (Fin Ops) initiatives across cloud workloads.
  • Support AWS (or other cloud platforms) operations and engineering needs.
  • Work with ROSA/container platforms for deployment, scaling, and reliability.
  • Recommend improvements in technology, architecture, and domain-specific reliability areas.
  • Manage and support large-scale systems operating at scale.
  • Reduce toil by identifying repetitive tasks and automating them.
  • Contribute code, read/interpret service repositories, and assist teams with engineering tasks as needed.

Required Skills & Experience

  • Strong background in SRE operations for distributed systems.
  • Proficiency in development/coding (Python, Go, shell scripting, or similar).
  • Ability to read/interpret codebases and build REST APIs.
  • Experience with Dynatrace/observability onboarding and ecosystem.
  • Deep knowledge of resiliency engineering and failover strategies.
  • Strong understanding of Fin Ops principles and cloud cost optimization.
  • Hands-on experience with AWS or any other cloud provider.
  • Experience with ROSA or Kubernetes-based container platforms.
  • Proven automation skills to eliminate operational toil.
  • Experience managing large-scale systems in production.
  • Capability to suggest architecture and domain improvements.
  • Strong analytical, troubleshooting, and collaboration skills.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary