Senior Engineering Manager, AI Infrastructure
Listed on 2026-07-11
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Senior Engineering Manager, AI Infrastructure
Seattle, WA
Our base salary range is $146,880 - $220,320, and in addition we have generous bonus plans to provide a competitive compensation package.
We are seeking a Senior Manager, AI Infrastructure to run the day-to-day operation of the systems that power our research. Reporting to the VP of Engineering, you will own the execution and reliability of our high-performance computing (HPC) environment which includes on-prem GPU clusters and the software orchestration layer that schedules workloads across a hybrid cloud environment. This is a hands-on operational leadership role: your mandate is to keep the platform fast, reliable, and well-utilized, and to deliver against the roadmap set with your PM counterpart.
Our ideal candidate is a:
- Systems Expert:
You have a deep, hands-on understanding of the Linux kernel, container runtimes, and distributed systems. You understand the performance implications of Infini Band topologies and NCCL optimizations. - Execution-Focused Leader:
You plan and deliver against near-term operational goals, keep reliability and researcher velocity high, and turn priorities set with leadership into shipped, dependable systems. - Pragmatic Operator:
You are comfortable making trade-offs between technical elegance and operational necessity. You triage and mitigate immediate risks, and know when to handle something yourself versus escalate.
Ai2 is a non-profit research institute at the forefront of open-source AI development. Unlike industry peers, our goal is to share our findings, data, code, and models with the global scientific community.
Why Ai2:- Open Science:
Your work directly enables the release of open models like OLMo, providing the broader research community with tools they can't get elsewhere. - Mission-Driven:
We prioritize scientific impact over profit margins. This allows us to focus on building the "right" infrastructure for long-term research goals. - Complexity at Scale:
You will manage some of the most dense and high-performance compute environments currently in operation.
Your Next Challenge:
- Cluster Operations:
Manage the availability, performance, and health of our dense on-prem GPU clusters. Coordinate with hardware vendors and internal teams to keep physical infrastructure meeting the demands of frontier model training. - Orchestration & Scheduling:
Operate and improve Beaker, our internal orchestration platform by optimizing resource allocation and driving high utilization across on-prem assets and elastic cloud resources (AWS/GCP). - Storage Operations:
Execute and continuously improve our storage environment, balancing high-throughput performance for active training against cost-effective durability for petascale research data. Contribute to the longer-term storage roadmap. - Resource Management:
Manage GPU compute allocation against budget. Track utilization, surface the data, and recommend when to burst to the cloud versus investing in on-prem capacity, escalating larger trade-offs as needed. - User Support & Velocity:
Serve as the technical bridge to our research teams. Ensure infrastructure is an accelerator, not a bottleneck, for a diverse set of research objectives. - Team Leadership:
Manage and grow a team of systems engineers, SREs, and software developers. Set the bar for operational rigor, engineering quality, and a collaborative culture, and keep the team unblocked and delivering.
What You'll Need:
- Experience:
12+ years in infrastructure, systems engineering, or HPC (or an advanced degree with 8+ years), including 2+ years supervising a small engineering team (5+). - Bachelor's degree in a related field: a relevant advanced degree may substitute for equivalent years of technical work experience.
- GPU/HPC Stack:
Direct experience operating large-scale NVIDIA GPU clusters and high-performance networking (Infini Band/RoCE). - Orchestration:
Strong background in Kubernetes, Slurm, or similar orchestration frameworks, particularly in hybrid-cloud configurations. - Storage:
Hands-on experience with distributed file systems (e.g., WEKA, Ceph, Lustre) and cloud storage integration at scale. - Software Development:
Proficient in designing and managing SDLC processes including sprint planning and technical design reviews. Proficient in Go or Python.
Physical Demands and Work Environment:
- Must be able to remain in a stationary position for long periods of time.
- The ability to communicate information and ideas so others will understand. Must be able to exchange accurate information in these situations.
- The ability to observe details at close range.
- Can work under deadlines.
A Little More About Ai2:
Ai2 is a Seattle based non-profit AI research institute founded in 2014 by the late Paul Allen. Our mission is building breakthrough AI to solve the world's biggest problems. We develop foundational AI research and innovation to deliver real-world impact through large-scale open models, data, robotics, conservation, and beyond.
In addition to Ai2's core mission, we also aim to…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).