Site Reliability Engineer (SRE
Listed on 2026-08-10
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, IT Support, Systems Engineer
Job Details
Job Location:
Atlanta, GA 30328
Salary Range: $ - $ Salary Site Reliability Engineer (SRE)
When you join Atlanticus, you become a member of a fast-growing, mission-focused company that is committed to aid in meeting the financial needs of middle-class Americans. With a culture of collaboration and a one-team mindset, we encourage entrepreneurial thinking to empower our customers toward financial well-being.
Atlanticus™ technology enables bank, retail, and healthcare partners to offer more inclusive financial services to everyday Americans through the use of proprietary analytics. We apply the experience gained and infrastructure built from servicing over 20 million customers and over $40 billion in consumer loans over more than 25 years of operating history to support lenders that originate a range of consumer loan products.
These products include retail and healthcare, private label credit and general-purpose credit cards marketed through our omnichannel platform, including retail point-of-sale, healthcare point-of-care, direct mail solicitation, digital marketing, and partnerships with third parties. Additionally, through our Auto Finance subsidiary, Atlanticus serves the individual needs of automotive dealers and automotive non-prime financial organizations with multiple financing and service programs.
Austin, TX – Situated in The Domain, a vibrant tech hub with park-like surroundings, top restaurants, and convenient parking, perfect for post-work socializing.
Atlanta, GA – Located in the Queen Building (King & Queen Towers, Sandy Springs), with easy access to I-285, GA-400, and a free shuttle to MARTA.
Work CultureWe foster a collaborative, innovative environment where everyone contributes to building something meaningful. You’ll be empowered to lead, grow, and make an impact.
The RoleWe are seeking a Site Reliability Engineer (SRE) to ensure the reliability, availability, performance, and operational excellence of our cloud-native applications running on AWS. This is a hands-on role responsible for monitoring and supporting production systems, automating operational tasks, managing deployments, and driving continuous improvements in system stability.
The ideal candidate has strong experience supporting Java-based applications running on Amazon EKS, a solid understanding of AWS infrastructure, and expertise with observability platforms such as Datadog and Splunk. This role requires participation in a 24x7 production support and on-call rotation, working closely with Development, Dev Ops, IT Ops, Database, Network, and Security teams to maintain highly available production services.
The successful candidate should be passionate about automation, troubleshooting complex production issues, improving application reliability, and leveraging AI-powered tools to enhance operational efficiency.
Key Responsibilities- Provide 24x7 production support through an on-call rotation to ensure application availability and rapid incident response.
- Continuously monitor production applications, infrastructure, and platform health using Datadog, Splunk, Cloud Watch, and other monitoring tools.
- Respond to production incidents, troubleshoot issues, and restore services while minimizing customer impact.
- Perform root cause analysis (RCA) and implement corrective actions to prevent recurring incidents.
- Deploy and support Java-based applications running on Docker and Amazon EKS using CI/CD pipelines.
- Execute production deployments, application releases, hotfixes, and rollbacks following change management processes.
- Monitor and manage scheduled application jobs, batch processes, and integrations to ensure successful execution.
- Troubleshoot Java application issues using logs, JVM metrics, thread dumps, heap dumps, and application performance metrics.
- Analyze application, infrastructure, and Kubernetes logs using Splunk and Datadog to identify performance bottlenecks and operational issues.
- Develop automation scripts using Python, Bash, or similar scripting languages to eliminate repetitive operational tasks.
- Build self-healing and automated operational processes to improve…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).