Site Reliability Engineer
Listed on 2026-08-24
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, IT Support
Site Reliability Engineer
Are you interested in working with the World's leading AI-powered Quality Engineering Company? Ready to advance your career, team up with global thought leaders across industries and make a difference every day? Join us at QualityAI! We are looking for a Site Reliability Engineer (SRE) to join our growing team in the United States!
Location:
Riverwoods, IL (Hybrid – 2 to 3 days/week onsite)
Position Overview:
We are seeking an experienced Site Reliability Engineer (SRE) with a strong background in AWS Cloud, monitoring/observability platforms, and automation. The ideal candidate will partner closely with application development teams to improve application reliability, resiliency, performance, and operational excellence across hybrid cloud environments. This role is ideal for engineers with hands-on experience in monitoring engineering using tools such as Datadog, Dynatrace, Grafana, Kibana, and expertise in AWS-based applications.
Must-Have
Skills:
- 6–12 years of professional experience as a Site Reliability Engineer (SRE)
- Strong hands-on experience with AWS Cloud applications and services (mandatory)
- Experience supporting hybrid environments (AWS Cloud and on-premises deployments)
- Strong Linux/Unix administration and shell scripting experience
- Experience with Systems Observability and Application Performance Monitoring (APM) tools, preferably Datadog (Dynatrace experience is also valuable)
- Experience building dashboards using Grafana and Kibana
- Experience in performance testing and the ability to translate functional and non-functional requirements into automated non-functional testing (NFT) solutions
- Strong understanding of application integration, high availability, resilience, and observability
- Experience with Dev Ops practices and CI/CD pipelines
- Strong programming skills in one or more of the following:
Python, Java, Shell Scripting (Unix/Linux) - Strong understanding of Software Development Life Cycle (SDLC)
Preferred Qualifications:
- Hands-on experience with Service Now (SNOW)
- Experience with container technologies such as Kubernetes and Open Shift
- Experience with Jenkins and CI/CD automation
- Experience with automation tools such as Ansible
- Strong working knowledge of JIRA
- Basic understanding of Release Management
- Understanding of Agile methodologies
- Experience with AWS Lambda services
- Knowledge of SQL, MySQL, and database concepts
Key Responsibilities:
- Partner with application development teams to improve application resiliency and reliability.
- Implement and maintain Service Level Objectives (SLOs), Service Level Indicators (SLIs), and operational best practices.
- Build end-to-end observability solutions using monitoring, logging, and tracing tools.
- Design and implement monitoring, alerting, dashboards, and health checks for production applications.
- Automate operational processes to reduce manual effort and improve system reliability.
- Develop and enhance capacity planning and performance management capabilities.
- Support disaster recovery (DR) planning and implementation for critical applications.
- Develop and support chaos engineering and resilience testing initiatives.
- Participate in production incident management and on-call support rotation.
- Collaborate with cross-functional teams to improve system availability, scalability, and operational efficiency.
Technical
Skills:
- Cloud: AWS (mandatory), Hybrid Cloud, On-Prem, Linux/Unix
- Monitoring & Observability:
Datadog (preferred), Dynatrace, Grafana, Kibana, ELK, APM tools - Programming:
Python, Java, Shell Scripting, Go (preferred), Ansible - Dev Ops:
Jenkins, CI/CD, Kubernetes, Open Shift - Tools: JIRA, Service Now (SNOW), Agile, Release Management
- Database: SQL, MySQL
Benefits:
Why QualityAI?
QualityAI is an AI-first quality engineering company helping enterprises deploy and scale complex systems with greater confidence. Operating across data, models, platforms, infrastructure, and operational environments, the company provides assurance and engineering expertise that helps organizations ensure systems perform reliably in real-world conditions. Formerly Qualitest, QualityAI supports global enterprises across regulated and technology-driven industries, combining…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).