Sr. Cloud Operations Reliability Engineer; SRE
Job in
Macon, Bibb County, Georgia, 31297, USA
Listed on 2026-07-24
Listing for:
NextGen Healthcare Information Systems LLC
Full Time
position Listed on 2026-07-24
Job specializations:
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Disaster Recovery IT
Job Description & How to Apply Below
Senior Cloud Operations Reliability Engineer
Responsibilities include:
- Drive operational excellence and strengthen the reliability posture of cloud-based services and supported platforms.
- Own critical reliability initiatives, establish observability and service health practices, and improve service availability, resiliency, and recovery.
- Lead incident response coordination and post-incident processes, including troubleshooting complex production issues, conducting root‑cause analysis, and driving remediation activities.
- Design and implement reliability‑focused automation, operational tooling, and runbooks to reduce manual toil, improve response consistency, and strengthen production readiness.
- Build observability solutions through comprehensive monitoring, logging, and alerting strategies; establish event correlation and escalation procedures.
- Conduct performance and capacity analysis, evaluate utilization trends, identify bottlenecks, and recommend improvements for scaling, performance, and capacity planning.
- Partner with development and engineering teams to evaluate deployment readiness, support deployment reliability improvements, and implement operational best practices that strengthen service reliability.
- Contribute to disaster recovery and business continuity planning, conduct operational readiness exercises, and keep recovery procedures up to date.
- Mentor team members, establish reliability standards, and create operational documentation, SOPs, and knowledge‑base materials.
- Support compliance, governance, and security initiatives, including access control, tagging, logging, and audit readiness.
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field – or an equivalent combination of education and experience.
- 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, Dev Ops, Infrastructure Operations, or related discipline.
- Extensive hands‑on experience supporting production cloud environments using Google Cloud Platform, AWS, or equivalent cloud service providers.
- Knowledge of monitoring, observability platforms, alerting strategies, incident response, root‑cause analysis, and production support in distributed or cloud‑native architectures.
- Experience with Infrastructure as Code (Terraform, Deployment Manager, Cloud Formation, etc.) and version control best practices.
- Strong background in incident management and post‑incident review processes with the ability to drive corrective actions.
- Experience with Kubernetes operations, containerization, and orchestration platforms.
- Experience with application performance monitoring and distributed tracing.
- Experience mentoring junior engineers or leading operational improvements initiatives.
- Mandatory certifications:
Google Cloud Associate Cloud Engineer, Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud Operations Engineer, or Google Cloud Professional Data Engineer; AWS Sys Ops Administrator or equivalent; advanced certifications in Kubernetes, Terraform, observability platforms, Dev Ops, Site Reliability Engineering, or ITIL.
- CI/CD practices, cloud governance, compliance frameworks, disaster recovery, and business continuity planning.
- Deep technical knowledge of Google Cloud Platform, AWS, or similar cloud providers; understanding of cloud‑native services, networking, security, and compute models.
- Familiarity with observability tools such as Grafana, Prometheus, Cloud Monitoring, Datadog, New Relic, etc.
- Proficiency in scripting languages (Python, Bash, Go) to develop automation solutions.
- Advanced ability to diagnose complex, multi‑layered infrastructure issues and coordinate timely recovery.
- Strong communication skills to translate complex technical findings into actionable recommendations and influence cross‑functional teams.
Next Gen Healthcare is an equal‑opportunity employer.
#J-18808-LjbffrTo View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×