×
Register Here to Apply for Jobs or Post Jobs. X

Sr. Cloud Operations Reliability Engineer; SRE

Job in Macon, Bibb County, Georgia, 31297, USA
Listing for: NextGen Healthcare Information Systems LLC
Full Time position
Listed on 2026-07-24
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Disaster Recovery IT
Salary/Wage Range or Industry Benchmark: 140000 - 190000 USD Yearly USD 140000.00 190000.00 YEAR
Job Description & How to Apply Below
Position: Sr. Cloud Operations Reliability Engineer (SRE)

Senior Cloud Operations Reliability Engineer

Responsibilities include:

  • Drive operational excellence and strengthen the reliability posture of cloud-based services and supported platforms.
  • Own critical reliability initiatives, establish observability and service health practices, and improve service availability, resiliency, and recovery.
  • Lead incident response coordination and post-incident processes, including troubleshooting complex production issues, conducting root‑cause analysis, and driving remediation activities.
  • Design and implement reliability‑focused automation, operational tooling, and runbooks to reduce manual toil, improve response consistency, and strengthen production readiness.
  • Build observability solutions through comprehensive monitoring, logging, and alerting strategies; establish event correlation and escalation procedures.
  • Conduct performance and capacity analysis, evaluate utilization trends, identify bottlenecks, and recommend improvements for scaling, performance, and capacity planning.
  • Partner with development and engineering teams to evaluate deployment readiness, support deployment reliability improvements, and implement operational best practices that strengthen service reliability.
  • Contribute to disaster recovery and business continuity planning, conduct operational readiness exercises, and keep recovery procedures up to date.
  • Mentor team members, establish reliability standards, and create operational documentation, SOPs, and knowledge‑base materials.
  • Support compliance, governance, and security initiatives, including access control, tagging, logging, and audit readiness.
Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field – or an equivalent combination of education and experience.
  • 10+ years of professional experience in Cloud Operations, Site Reliability Engineering, Dev Ops, Infrastructure Operations, or related discipline.
  • Extensive hands‑on experience supporting production cloud environments using Google Cloud Platform, AWS, or equivalent cloud service providers.
  • Knowledge of monitoring, observability platforms, alerting strategies, incident response, root‑cause analysis, and production support in distributed or cloud‑native architectures.
  • Experience with Infrastructure as Code (Terraform, Deployment Manager, Cloud Formation, etc.) and version control best practices.
  • Strong background in incident management and post‑incident review processes with the ability to drive corrective actions.
  • Experience with Kubernetes operations, containerization, and orchestration platforms.
  • Experience with application performance monitoring and distributed tracing.
  • Experience mentoring junior engineers or leading operational improvements initiatives.
  • Mandatory certifications:
    Google Cloud Associate Cloud Engineer, Google Cloud Professional Cloud Architect, Google Cloud Professional Cloud Operations Engineer, or Google Cloud Professional Data Engineer; AWS Sys Ops Administrator or equivalent; advanced certifications in Kubernetes, Terraform, observability platforms, Dev Ops, Site Reliability Engineering, or ITIL.
Knowledge, Skills & Abilities
  • CI/CD practices, cloud governance, compliance frameworks, disaster recovery, and business continuity planning.
  • Deep technical knowledge of Google Cloud Platform, AWS, or similar cloud providers; understanding of cloud‑native services, networking, security, and compute models.
  • Familiarity with observability tools such as Grafana, Prometheus, Cloud Monitoring, Datadog, New Relic, etc.
  • Proficiency in scripting languages (Python, Bash, Go) to develop automation solutions.
  • Advanced ability to diagnose complex, multi‑layered infrastructure issues and coordinate timely recovery.
  • Strong communication skills to translate complex technical findings into actionable recommendations and influence cross‑functional teams.

Next Gen Healthcare is an equal‑opportunity employer.

#J-18808-Ljbffr
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary