Senior Site Reliability Engineer
Listed on 2026-07-13
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, AWS
QGenda is redefining healthcare workforce management everywhere care is delivered. We're on a mission to empower the healthcare industry to better onboard, deploy, and manage their workforce. Over 4,500 healthcare organizations have trusted us to help them make strategic workforce decisions through our unified software platform. With more than 800 employees across the US, we are united in our vision and culture to make a difference for our customers and enjoy the day-to-day.
At QGenda, we value our employees and their contributions toward the success of the business. We strive to create a dynamic work environment that fosters growth, innovation, and collaboration, where employees can be proud of the work they do and the impact it has on the healthcare industry. QGenda is headquartered in Atlanta. To learn more about QGenda, visit us at or follow us on Instagram or Linked In.
As a Senior Site Reliability Engineer, you will work with our Infrastructure and Product Development Teams to increase the scalability, reliability, and performance of our systems and services. You will build and extend existing automation for configuration and monitoring of our AWS hosted applications, evaluate new AWS services and tools, and focus on platform health and monitoring to deliver the best experience for our customers.
This role is hybrid with one required day in our Buckhead (Atlanta, Georgia) or Uniontown, Ohio office depending on your current location.
- Design, implement, and manage scalable systems that ensure high availability, fault tolerance, and optimal performance.
- Continuously monitor and enhance system health and performance through data analysis and metrics.
- Develop and advocate for automation tools to eliminate repetitive manual processes and improve efficiency.
- Build and enhance CI/CD pipelines to streamline software delivery and deployments.
- Participate in on‑call rotation to respond to incidents, troubleshoot problems, and minimize downtime.
- Conduct root cause analyses and implement permanent solutions to recurring issues.
- Manage our cloud‑based infrastructure environment in AWS.
- Optimize costs and resources while maintaining robust and scalable systems.
- Serve as a technical advisor to engineering teams on infrastructure and operations best practices.
- Actively contribute to fostering an SRE culture within the organization by promoting observability, retrospectives, and continuous improvement.
- Curiosity‑driven mindset with a desire to continuously learn and improve systems
- Strong sense of ownership — you see problems through to resolution, not just escalation
- Comfortable navigating ambiguity and making pragmatic tradeoffs under pressure
- Availability for off‑hours deployment and upgrades of production systems during release and maintenance windows
- Strong problem‑solving skills and ability to work effectively under pressure.
- Excellent communication skills for cross‑functional collaboration as well as documentation creation.
- B.S. in Computer Science, Computer Information Systems, or Computer Engineering from a major U.S. university or equivalent industry experience
- 7+ years of experience as a Dev Ops, SRE or Systems Engineer
- Advanced proficiency with at least one scripting or programming language
- Experience with Docker and container orchestration tools such as AWS ECS and EKS/Kubernetes
- Hands‑on experience building infrastructure and supporting applications in AWS using services such as Lambda, EC2, ECS, S3, SNS, SQS, RDS, Redshift, and Elasticache
- Strong understanding of networking and DNS
- Strong experience with Terraform for infrastructure provisioning and module development, along with configuration management and infrastructure as code (IaC) practices
- Firm understanding and experience with Agile and Scrum SDLC processes
- Experience using distributed version control system (Git preferred) to check‑in code, branching, merging, pull request, code review, etc
- Knowledge of CI/CD best practices and tools such as AWS Code Build,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).