Software Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Job in
Charlotte, Mecklenburg County, North Carolina, 28245, USA
Listed on 2026-08-19
Listing for:
Judge Group, Inc.
Full Time
position Listed on 2026-08-19
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below
Senior Site Reliability Engineer (SRE)
We are not accepting C2C or 1099 arrangements.
Location: Charlotte, NC (Preferred) or Chandler, AZ
Employment Type: Contingent / Contract Assignment
About the Role
We are looking for a Senior Site Reliability Engineer (SRE) to help drive the reliability, scalability, and security of enterprise platforms across Windows, Linux, and cloud-native environments. In this role, you will support the transformation from traditional application support models to modern platform engineering practices. You will leverage expertise in Google Cloud Platform (Google Cloud Platform), automation, containerization, and infrastructure engineering to build resilient systems that enable business-critical applications at scale.
As a member of the Site Reliability Engineering team, you will collaborate with software engineers, infrastructure teams, and security partners to improve platform availability, operational efficiency, and cloud adoption.
Responsibilities
Platform Reliability and Cloud Engineering
- Design, implement, and maintain highly available, scalable, and secure production systems across Windows, Linux, and Google Cloud Platform environments.
- Build and support containerized platforms using Kubernetes (GKE) and Docker.
- Develop and manage infrastructure through Infrastructure-as-Code (IaC) tools including Terraform and Ansible.
- Improve platform performance, reliability, and capacity through proactive engineering and optimization.
- Create automation solutions to reduce operational overhead and improve incident response efficiency.
- Develop monitoring, alerting, and observability capabilities using SLIs, SLOs, Prometheus, Grafana, and Google Cloud Operations Suite.
- Implement telemetry and performance metrics across hybrid and cloud environments.
- Lead incident response efforts, perform root cause analyses, and facilitate post-incident reviews.
- Design and implement self-healing systems and automated remediation workflows.
- Drive continuous improvement initiatives to enhance system reliability and operational excellence.
- Partner with Information Security teams to implement security best practices, vulnerability management, and compliance requirements.
- Integrate security controls into cloud platforms, infrastructure, and CI/CD pipelines.
- Support identity management, encryption, access controls, and policy enforcement across enterprise environments.
- Work closely with software developers, application owners, and infrastructure engineers to build reliable cloud-native solutions.
- Develop and maintain technical documentation, operational procedures, and runbooks.
- Serve as a trusted technical advisor on platform reliability and operational best practices.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or equivalent practical experience.
- 5+ years of experience in Software Engineering, Site Reliability Engineering, Systems Engineering, or related technical roles.
- 3+ years of hands-on experience supporting production Windows and/or Linux environments.
- Experience administering and troubleshooting large-scale production systems.
- Experience with infrastructure automation and scripting using Power Shell, Python, Shell, or similar languages.
- Experience with Google Cloud Platform (Google Cloud Platform), including GKE, IAM, Cloud Functions, Cloud Monitoring, and related services.
- Experience with container orchestration technologies, including Kubernetes and Docker.
- Experience with Infrastructure-as-Code tools such as Terraform and Ansible.
- Strong understanding of Linux system administration and hybrid cloud architectures.
- Knowledge of Active Directory, DNS, DHCP, and Windows security concepts.
- Experience implementing CI/CD pipelines using tools such as Git Lab CI, Jenkins, or similar platforms.
- Familiarity with ITIL practices, change management processes, and incident management frameworks.
- Experience with Service Now, load balancers, certificate…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×