Azure Cloud Engineer
Listed on 2026-08-22
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer, Azure
Location:
Charlotte, NC;
Chandler, AZ; or Dallas, TX
Hybrid: 3 days onsite, 2 days remote
Duration: 12-month contract with strong potential for extension
Our client, a leading U.S.
-based financial services institution, is seeking an experienced Azure Cloud Platform Engineer to support, maintain, and enhance large-scale Azure platform and application hosting environments. This role will focus on cloud operations, platform stability, application hosting, observability, automation, and continuous improvement across cloud and on-premises environments.
The ideal candidate brings strong hands-on experience with Azure, Kubernetes, Linux, infrastructure automation, monitoring and observability, and production platform support. This individual will also help drive automation, self-service capabilities, toil reduction, and operational efficiency while partnering with engineering teams, technical stakeholders, vendors, and senior leadership.
In This Role, You Will:- Support and maintain large-scale Azure platform and application hosting environments, including Azure Kubernetes Service (AKS) and Azure App Service Environment (ASE).
- Lead or participate in complex, broad-impact initiatives involving cloud infrastructure, platform services, systems operations, and application hosting.
- Provide L2/L3 operational support for production Azure environments and troubleshoot complex platform and infrastructure issues.
- Lead initiatives to improve observability, monitoring, alerting, platform stability, automation, and operational efficiency.
- Identify operational gaps, inefficiencies, and sources of manual effort and drive continuous improvement and toil-reduction initiatives.
- Perform incident management, root cause analysis, change management, capacity planning, platform maintenance, and day-to-day operational oversight.
- Support and troubleshoot Azure networking, identity, security, compute, and platform services.
- Work with engineering teams on technical changes, platform enhancements, and solution design while ensuring adherence to technical controls, governance standards, and enterprise architecture guidelines.
- Develop and maintain automation and Infrastructure-as-Code (IaC) to streamline provisioning, configuration, operational support, and self-service capabilities.
- Support containerized workloads and Kubernetes platform operations.
- Review and analyze technical and operational challenges, evaluate alternatives, and implement platform improvements and corrective actions.
- Collaborate with technical peers, cross-functional teams, vendors, and mid-to-senior level stakeholders to resolve issues and maintain high levels of systems and infrastructure availability.
- Participate in Change and Release Management processes to minimize service disruption and maintain platform stability.
- Maintain technical documentation, operational procedures, and platform support standards.
- Provide clear communication and status updates regarding incidents, platform health, operational risks, and technical initiatives.
- Help mentor and provide technical guidance to other engineers as needed.
- 7+ years of experience supporting public cloud and on-premises technologies.
- 4+ years of experience supporting Linux environments.
- 3+ years of hands-on experience with Azure Kubernetes Service (AKS), Azure App Service Environment (ASE), and Azure Landing Zones.
- 3+ years of experience with Terraform and Git Hub.
- 3+ years of experience with Python or comparable scripting/automation experience.
- 2+ years of experience with Grafana, KQL, and JQL.
- Experience supporting production-scale cloud and application hosting environments.
- Strong experience with incident management, change management, platform troubleshooting, and operational support.
- Strong understanding of cloud infrastructure, platform stability, observability, monitoring, and automation.
- Experience reducing manual operational effort through automation and self-service capabilities.
- 1+ year of experience with LLMs, prompting, conversational AI, agentic AI, and/or MCP.
- Strong verbal, written, analytical, and interpersonal communication skills.
- Ability to work independently while collaborating effectively across engineering, operations, and business teams.
- 3+ years of experience supporting shared hosting environments.
- 3+ years of experience in Site Reliability Engineering (SRE).
- Experience supporting Open Shift platforms.
- Strong experience with Service Now for incident and change management.
- 1+ year of experience with Jira and Service Now.
- Experience with Azure Monitor, Prometheus, or other monitoring and observability platforms.
- Experience with Docker, Helm, or other containerization and orchestration technologies.
- Experience with Azure networking, identity, security, and enterprise cloud governance.
- Experience with Azure Virtual Machines, Azure Virtual Networks, NSGs, Private Endpoints, and Entra .
- Experience with Infrastructure-as-Code and automation tools such as Terraform, ARM, Bicep, or Ansible.
- Familiarity…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).