Director of Infrastructure Operations
Listed on 2026-07-14
-
IT/Tech
Systems Administrator, Cloud Computing: Infrastructure & Operations, IT Project Manager, IT Infrastructure
Position Overview
The Director will manage a team of Engineers and Technicians responsible for maintaining a 99.99% uptime for data center and cloud infrastructure and related IT business services. The Director will be responsible for maintaining, monitoring, configuring, and supporting all US infrastructure. The Director will be the escalation point for internal customers and break/fix issues, with ultimate responsibility for incident tracking, resolution, and reporting.
The Director will be responsible for continuous platform improvements through the introduction of technologies that achieve a positive ROI and/or an improvement to the customer experience.
The Director is responsible for 24×7 management and production operations for all infrastructure systems under the management. This includes all internally managed and outsourced engineering of the Enterprise Infrastructure Services including Windows, Unix, Linux, physical and virtual storage (SAN, NAS services and backup/recovery systems), virtualization platforms, etc. This includes both on-prem, datacenter and cloud infrastructure (primarily AWS). Responsibilities include operational management (development through production);
IT governance of these systems; global monitoring standards and observability; patch management currency; and budgetary planning and management (capital and operational).
- Drive a cloud-first, automation-first operations strategy, leveraging AWS-native services, Infrastructure as Code (IaC), and event-driven automation to minimize manual intervention and improve operational efficiency.
- Lead the team in designing, implementing, and maintaining automated operational workflows using tools such as CI/CD pipelines, scripting (Python), and cloud-native automation frameworks.
- Establish and enforce observability standards across all platforms (on‑prem and AWS), including logging, metrics, tracing, and AIOps-driven insights for predictive operations.
- Manage lifecycle of infrastructure assets across hybrid environments, including servers, storage, network devices, and cloud resources, with emphasis on cloud resource governance and optimization.
- Maintain 99.99% SLA uptime of production and production‑supporting infrastructure.
- Responsible for incident management and issue‑resolution processes, including establishment of monitoring and management of alerts, escalations, and notifications.
- Develop, document, and implement operating policies, methods, processes, measures, tools, MOPs, SOPs, and training programs.
- Manage MSPs and ensure SLA/KPI attainment and SOW compliance.
- Manage annual expense budgets.
- Support capacity planning for backup assets, power, network bandwidth, circuits, etc.
- Orchestrate and lead annual DR testing.
- Implement and enforce change management process.
- Advocate and enforce adherence to compliance (SOX, PCI, etc.), as well as technological architectural designs and standards.
- Responsible for problem resolution and tracking of events using ticketing system.
- Lead the team ensuring stellar customer service and compliance with customer Service Level Agreements (SLA), and IT business service reliability, availability, preventive maintenance practices, and readiness of network components. Active ITSM focus and continuous improvement by measuring and managing the team against KPIs.
- Ensure that operation management policies, standards and procedures are properly implemented, enhanced, and maintained, including training, incident reporting and resolution, issue management, metrics collection, performance reporting, trend analysis, and Continuity of Operations (COOP).
- Responsible for developing and measuring the effectiveness of L3 support structures and approaches inclusive of measurement and improvement programs.
- Ensures that L1/L2 service levels by supplier are maintained and continually improved upon and ensures all support and remediation processes and procedures are stringently followed.
- Key stakeholder and executive sponsor of Problem Management and Root Cause Analysis processes.
- Thought leader, evangelist and change agent for all tools and platforms that support and influence the Core Operations Engineering…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).