Lead Cloud Engineer
Listed on 2026-07-18
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Azure, Systems Engineer
Service Reliability and Monitoring
- Monitor availability, performance, and health of Azure-hosted internal applications using tools such as Azure Monitor, Application Insights, Open Telemetry and Log Analytics
- Participate in incident response workflows including triage, escalation, and post-incident review with a focus on reducing mean time to recovery (MTTR)
- Maintain and refine alerting thresholds and dashboards to surface actionable signals over noise
- Contribute to SLO/SLA tracking and reporting for supported internal platforms
- Support CI/CD pipeline operations using Azure Dev Ops, including pipeline monitoring, failure triage, and basic remediation
- Assist in deploying updates and configuration changes to internal Azure-based applications following change management practices
- Identify opportunities to automate repetitive operational tasks using scripting (Power Shell, Python, or Bash)
- Maintain infrastructure-as-code contributions under guidance (Bicep, ARM templates, or terraform)
- Experience with Git based repos, pull requests
- Provide operational support for internal AI-powered and agent-based tools running on Azure, including availability monitoring and incident coordination
- Work alongside engineering teams to understand deployment dependencies for agentic workloads and flag reliability risks early
- Assist with environment validation and regression checks following updates to AI-integrated services
We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances.
If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy:
- 5+ years of experience in an IT operations, Dev Ops, or systems support role (internship or co-op experience considered)
- Foundational understanding of cloud infrastructure concepts with hands-on or coursework exposure to Microsoft Azure
- Familiarity with CI/CD concepts and at least one pipeline toolset (Azure Dev Ops preferred)
- Basic scripting ability in Power Shell, Python, or Bash
- Strong written communication skills -- this role documents as much as it resolves
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).