Lead Infrastructure Engineer - Messaging & Collaboration Operations
Listed on 2026-08-27
-
IT/Tech
SRE/Site Reliability, Systems Engineer
Lead Infrastructure Engineer
The Lead Infrastructure Engineer is responsible for the operational health, reliability, and continuous improvement of enterprise messaging and collaboration platforms. This role provides technical and operational leadership for Microsoft-based communication and productivity services, ensuring stable, secure, and highly available experiences for employees across the organization.
The successful candidate will lead major incident response activities, drive problem management initiatives, improve service reliability through automation and monitoring, and partner closely with Engineering, Product, Security, Infrastructure, and Support teams. This individual serves as a technical leader, operational strategist, and trusted advisor focused on operational excellence and service resiliency.
In this role, you will have the following responsibilities:
Enterprise Messaging & Collaboration Operations- Lead operational support for Microsoft 365, Exchange Online, Microsoft Teams, SharePoint Online, One Drive, Microsoft Copilot, and related collaboration technologies
- Ensure platform availability, performance, reliability, and operational readiness across enterprise environments.
- Develop and maintain operational standards, runbooks, knowledge articles, and support procedures
- Partner with engineering teams to transition new services, features, and platform changes into production support
- Serve as an escalation point for high-severity production incidents and complex service degradations.
- Lead Major Incident Management activities from detection through restoration and executive communication.
- Coordinate technical response teams during service disruptions and drive clear ownership of next actions.
- Conduct post-incident reviews and ensure corrective actions are documented, assigned, and tracked to closure.
- Establish processes that reduce Mean Time to Detect and Mean Time to Restore.
- Drive root cause analysis for recurring issues, chronic service degradations, and systemic operational risks.
- Lead problem management activities to eliminate repeat incidents and improve platform resiliency.
- Track operational trends and develop strategies to improve service health, supportability, and user experience.
- Partner with engineering teams to improve platform architecture and operational readiness.
- Design, optimize, and maintain monitoring strategies across critical collaboration services.
- Develop actionable dashboards, alerts, trend reports, and service health views using tools such as Splunk.
- Analyze logs and operational telemetry to identify service anomalies and emerging risks.
- Reduce alert fatigue through tuning, suppression logic, event correlation, and automation.
- Lead automation initiatives that improve operational efficiency, reduce manual effort, and strengthen auditability.
- Develop scripts, workflows, and tooling to automate recurring operational tasks and validation activities.
- Implement self-healing or guided remediation capabilities where appropriate.
- Promote continuous improvement focused on reliability, scalability, efficiency, and customer experience.
- Ensure compliance with IT Service Management processes for incidents, requests, problems, changes, and knowledge.
- Manage operational work through Service Now and maintain high-quality documentation and ticket hygiene.
- Support service-level objectives, service-level agreements, operational metrics, audits, and risk assessments.
- Contribute to governance frameworks supporting enterprise collaboration services.
- Provide clear, concise, and timely communications during operational events and service-impacting issues.
- Create executive-level summaries of incidents, trends, risks, and service performance.
- Partner with Product, Engineering, Security, Infrastructure, Support, and business stakeholders to establish priorities.
- Build trusted relationships across technical and business teams.
Required Qualifications:
- 5+ years of Technology Infrastructure Engineering and Solutions experience, or equivalent demonstrated through one or a combination of the following: work experience, training, military experience, education
- 5+ years of Engineering experience in Messaging and Collaboration applications, including on-premise and M365 within a large enterprise
- 5+ years of Power Shell scripting experience
Desired
Qualifications:
- MS-900 (M365 Fundamentals) or MS-700 (Teams Communications Administration)
- AZ-900 (Azure Fundamentals)
- Experience supporting hybrid messaging environments
- Experience in regulated enterprise environments
- Experience implementing Site Reliability Engineering principles
Leadership Expectations:
- Lead through influence, ownership, and technical expertise.
- Mentor engineers and operations analysts through knowledge sharing and practical coaching.
- Foster a culture of accountability,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).