Cloud Engineer III
Listed on 2026-06-25
-
IT/Tech
Windows Server, Systems Engineer, Azure, Disaster Recovery IT
The System Engineer Level 3 (Windows & Azure) is a senior-level technical role responsible for leading advanced system engineering, cloud operations, and infrastructure modernization across corporate, retail, and cloud environments. This position owns the resolution of high-impact and complex incidents, provides technical leadership for Windows, Microsoft 365, and Azure platforms, and ensures systems are secure, scalable, and operationally mature. The role drives improvements in identity and access management, endpoint management, patching, automation, monitoring, and disaster recovery while operating within established standards and change control processes.
This role also supports Fin Ops accountability by guiding resource governance, optimizing cloud consumption, and improving cost visibility through tagging discipline, usage reviews, and proactive cost optimization actions.
- Lead Tier 3 troubleshooting and resolution for complex Windows, Microsoft 365, and Azure incidents impacting business‑critical services.
- Serve as the senior technical escalation point for Windows Server systems, enterprise authentication, and access‑related issues across corporate and retail environments.
- Engineer, maintain, and improve core infrastructure services including Active Directory, Group Policy, identity integration, and endpoint management platforms.
- Lead Microsoft 365 administration and service health response activities, including advanced troubleshooting and coordination with vendors when required.
- Design and implement Azure solutions supporting compute, monitoring, backup, recovery, and operational stability using approved standards and best practices.
- Own and improve patching strategy and compliance outcomes for Windows endpoints and servers, including remediation workflows and reporting.
- Develop and maintain automation using Power Shell and platform tooling to reduce manual work, improve consistency, and increase reliability of operational processes.
- Lead endpoint management improvements including Autopilot/Intune strategy, device compliance, policy enforcement, and lifecycle governance.
- Establish and improve monitoring and alerting standards to support proactive detection, faster response, and reduced downtime across systems and cloud services.
- Support disaster recovery planning and execution, including backup validation, recovery testing, and documentation of restoration procedures.
- Lead change planning and execution for infrastructure improvements, ensuring risk assessment, change validation, rollback planning, and accurate documentation.
- Support Fin Ops accountability by enforcing Azure governance standards, improving tagging compliance, validating resource ownership, and supporting cost allocation reporting.
- Identify and drive cost optimization initiatives such as rightsizing, lifecycle management, cleanup of unused resources, and reduction of unnecessary consumption.
- Mentor Level 1 and Level 2 engineers through technical coaching, escalation support, and knowledge transfer to improve team capability and consistency.
- Participate in on‑call rotation and provide leadership during major incidents, including coordination, communication, and root cause documentation.
- Expert‑level knowledge of Windows OS and Windows Server administration in an enterprise environment.
- Strong experience administering Active Directory, Group Policy, and enterprise identity services, including troubleshooting complex authentication and access issues.
- Strong Microsoft 365 administration and troubleshooting experience across core services and integrations.
- Strong Azure engineering experience including compute operations, monitoring, backup/recovery, security fundamentals, and operational governance.
- Proven ability to lead technical design decisions, develop standards, and deliver improvements that increase reliability and supportability.
- Strong Power Shell scripting and automation capabilities with a focus on operational efficiency and consistency.
- Strong understanding of patching strategy, compliance reporting, and remediation workflows for endpoints and servers.
- Strong understanding of…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).