Operations Support System Engineer
Listed on 2026-08-28
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Azure, SRE/Site Reliability
Location: Chandler, AZ / Dallas, TX / Charlotte, NC — Hybrid, 3+ days onsite
Duration: 12 Months
Rate: $46.53/hr.
This 12-month contract opportunity is seeking a Systems Operations Engineer to support enterprise-scale Azure platform and application hosting environments across Chandler, AZ, Dallas, TX, or Charlotte, NC. This engineer will work extensively with Azure Kubernetes Service (AKS), Azure App Service Environment (ASE), Azure Landing Zones, Linux, Terraform, Git Hub, Python, and observability technologies. The position is W2 only and requires 3+ days per week onsite at an approved location.
This is an opportunity to help shape next-generation cloud operations rather than simply maintain existing infrastructure. You’ll tackle large-scale Azure environments, improve platform reliability and observability, automate repetitive operational work, and build self-service capabilities that reduce engineering toil. The role offers significant technical ownership, exposure to modern AI technologies such as LLMs and Agentic AI, and collaboration with engineering teams and senior leadership on high-impact cloud initiatives.
Contract Duration: 12 Months
Required Skills & Experience- 7+ years of experience with public cloud and on-premises technologies
- 4+ years of experience supporting Linux environments
- 3+ years of experience with Azure Kubernetes Service (AKS), Azure App Service Environment (ASE), and Azure Landing Zones
- 3+ years of experience with Terraform and Git Hub
- 3+ years of Python experience
- 1+ years of experience with LLMs, prompting and Conversational AI, Agentic AI, and Model Context Protocol (MCP)
- 2+ years of experience with Grafana, KQL, and JQL
- Experience supporting production enterprise Azure environments
- Strong incident response, troubleshooting, and operational support experience
- Strong verbal, written, and interpersonal communication skills
- Ability to independently support platform stability, incident management, change management, automation, documentation, and capacity planning
- Ability to participate in an on-call rotation and lead escalations and major incident support
- 3+ years of experience supporting shared hosting environments
- 3+ years of Site Reliability Engineering experience
- Experience with Open Shift
- Experience with incident and change management processes and Service Now
- 1+ years of experience with Jira and Service Now
- Experience with Azure networking, including VNets, NSGs, and Private Endpoints
- Experience with Microsoft Entra
- Familiarity with Docker, Helm, and service mesh technologies
- Experience with Prometheus and Azure Monitor
- Experience with Bash and/or Power Shell
- Experience with ARM, Bicep, and/or Ansible
- Understanding of cloud governance, policy enforcement, security best practices, and cost management
- Experience mentoring junior engineers or serving in a technical lead capacity
- Azure or Kubernetes certifications such as Azure Solutions Architect Expert, Azure Dev Ops Engineer Expert, Azure Administrator Associate, or CKA
- Support and maintain large-scale Azure platforms, including AKS and ASE environments
- Provide L2/L3 operational support and lead escalations and major incident response
- Improve observability, monitoring, alerting, platform stability, and overall system availability
- Develop automation and Infrastructure-as-Code solutions to reduce manual intervention and operational toil
- Support application migrations and enterprise cloud platform initiatives
- Conduct root cause analysis and implement corrective and preventative measures
- Participate in incident, change, and release management processes
- Identify operational gaps and drive continuous improvement initiatives
- Collaborate with engineering teams, vendors, cross-functional partners, and senior leadership
- Maintain technical documentation and support capacity planning and day-to-day operational oversight
- Grafana / KQL / JQL / Observability & Monitoring
- LLMs / Conversational AI / Agentic AI / MCP
- Service Now / Jira
- 70% Hands On — Azure platform support, troubleshooting, automation, observability, incident response, and infrastructure engineering
- 10% Management Duties — Leading escalations, coordinating major incidents, planning changes, and driving operational improvements
- 20% Team Collaboration — Partnering with engineering teams, cross-functional stakeholders, vendors, and senior leadership
Note: Daily responsibility percentages were not provided and are estimated based on the supplied responsibilities.
#J-18808-Ljbffr(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).