Technical Lead Manager, Infrastructure Hardware; Server and Network Systems
Listed on 2026-08-24
-
IT/Tech
Systems Engineer, IT Infrastructure
Technical Lead Manager, Infrastructure Hardware (Server and Network Systems)
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups.
OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
As a Technical Lead Manager, Infrastructure Hardware (Server and Network Systems) on the Cluster Architecture Team, you will provide technical leadership and drive end-to-end execution of server and network platform programs—including new product introductions (NPIs)—across Cerebras CS-3–based AI clusters. You will lead programs from requirements and technical trade-offs through vendor selection, lab bring-up, qualification, and production rollout. You will be the technical execution owner across OEM/ODM partners, component vendors, internal software/runtime teams and architects, validation/QA, and deployment/operations.
This role combines technical leadership with hands-on execution. You must understand server, network, and system-level design deeply enough to lead technical reviews, make and document practical trade-offs, and unblock execution—while partnering closely with Compute, Server Platform, and Network Architects on broader architectural direction. You will also build shared understanding with our rack/elevations and physical datacenter design partners so that server and network changes land smoothly in real deployments, without owning physical datacenter design.
Responsibilities
- Own end-to-end technical execution for server systems and network equipment in Cerebras clusters, including NPIs, platform refreshes, and major component or configuration changes.
- Drive requirements gathering and technical trade-off decisions, converting inputs into executable plans with clear milestones, readiness gates, and cross-functional deliverables.
- Represent Cluster Architecture in executive reviews, OKR cycles, and leadership or customer forums as needed.
- Build and manage integrated execution plans across vendors and internal teams, tracking dependencies, critical paths, and risks.
- Lead OEM/ODM, switch-vendor, and component-vendor engagements, including RFI/RFP activities, technical evaluations, samples, escalations, and roadmap alignment.
- Partner with Compute, Server Platform, and Network Architects to translate architectural direction into platform decisions, qualification plans, acceptance criteria, and rollout strategies.
- Lead NPI execution, qualification, and release readiness, including lab/staging validation, regression tracking, issue resolution, and go/no-go decisions.
- Step into execution gaps as needed to drive technical issues to closure and keep server and network programs moving.
- Own risk and change management into production, including versioning, rollout sequencing, and stakeholder communication.
- Ensure operational readiness with deployment and fleet teams and maintain alignment with rack and physical datacenter owners on power, cooling, space, and cabling constraints.
Skills and Qualifications
- B.S. or M.S. in Computer Science, Electrical/Computer Engineering, or equivalent experience.
- 8+ years in technical leadership, systems engineering, or technical program leadership for server, network, or infrastructure platforms from concept through production.
- Experience technically leading complex server and/or datacenter network programs across OEM/ODMs, switch vendors, component suppliers, and internal engineering teams.
- Strong knowledge of server architecture—including CPU/NUMA, memory bandwidth, PCIe, NIC, and storage I/O—and networking fundamentals including leaf-spine fabrics, switch platforms, optics, and high-performance interconnects.
- Fam…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).