Director, Network and Data Center Services
Listed on 2026-08-22
-
IT/Tech
Network Engineer, Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Working at Yale means contributing to a better tomorrow. Whether you are a current resident of our New Haven-based community, eligible for opportunities through the New Haven Hiring Initiative, or a newcomer, interested in exploring all that Yale has to offer, your talents and contributions are welcome. Discover your opportunities at Yale!
Yale ITS is seeking a senior technical leader to lead Network Engineering, Network Operations, and Data Center Services—responsible for the design, reliability, and evolution of Yale’s campus-wide communication and hosting infrastructure.
This role combines engineering authority, operational accountability, and transformation leadership, with responsibility for:
- A software-defined campus network
- A mandate to improve network operations through automation, observability, and discipline
- Two capacity‑constrained, at‑risk on‑prem data centers, a nearby co-lo and a cloud interconnection co-lo
This leader will drive end‑to‑end ownership of core network infrastructure, ensuring Yale’s network and data center platforms are resilient, scalable, and operationally excellent.
Key Responsibilities Network Strategy, Engineering & Transformation- Manage the architecture, engineering, and lifecycle management of Yale’s software-defined campus network
- Define standards for: performance, resiliency, segmentation, and security
- Drive modernization toward automated, code-defined network operations and integrated observability and telemetry
- Own 24x7 network operations, including availability, performance, and incident response
- Lead transformation of network operations tooling transitioning from Logic Monitor / Cisco-native tools → next-gen platforms (e.g., Datadog)
- Establish disciplined operational practices around monitoring, alerting, escalation, and incident management
- Improve key metrics and publish live network data for ITS leadership: MTTR, uptime, latency, incident recurrence
- Own the end-to-end roadmap for Yale’s data centers, addressing:
Capacity constraints;
Operational risk;
Lack of appropriate resiliency and DR capability - Define and execute a multi-year strategy across:
On-prem, colo, and cloud integration;
Improved resiliency;
Scalable capacity to support research and enterprise growth
- Embed resiliency by design across network and data center services by eliminating single points of failure and ensure network failover and recovery capabilities are validated
- Drive alignment with Disaster Recovery (RTO/RPO) targets
Champion automation across infrastructure operations:
Network automation;
Data center workflows;
Configuration and lifecycle management
- Build a unified observability model across networks, data centers, and facilities/campus operations
- Ensure monitoring tools provide end to end network visibility with actionable alerts integrated with incident management
- Lead and develop network engineering, operations, and data center teams while building a culture of ownership and accountability, engineering rigor, and an automation-first mindset
- Improve team capabilities in software-defined networking, automation and observability, and modern operations practices
- Work closely with core infrastructure engineering, Tech Ops, and Information security
- Navigate a federated organization, influencing without direct authority
- Align network and data center strategy with broader IT Stability initiatives, faculty and research needs.
- Proven leadership of large scale network and infrastructure environments (campus or enterprise, with hybrid cloud environments)
- Deep technical expertise in:
Enterprise networking (LAN/WAN, SDN, campus networks, large Cisco implementation);
Data center infrastructure (power, cooling, compute integration);
Observability and monitoring platforms - Experience leading 24x7 network operations environments and major transformation initiatives (tooling, architecture, operating model)
- Strong understanding of resiliency engineering, DR design and ITSM…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).