Data Center Operations Engineer
Listed on 2026-07-18
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, Hardware Engineer
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.
We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.
About the RoleCrusoe Cloud operates GPU infrastructure across six production sites globally, with a fleet that spans Super Micro, HPE, and next-generation ODM platforms as we scale. We're looking for a Staff Data Center Operations Engineer to serve as the senior technical operations resource for the Site Ops org — based at our Denver headquarters, with cross-site scope and travel authority across our full portfolio.
This role is the bridge between Crusoe's distributed site teams and our OEM and ODM hardware partners. You'll own platform-level escalations that exceed site-level capability, drive hardware decisions at the org level, and serve as Site Ops' technical presence at headquarters — visible to engineering, procurement, and leadership in a way that a field-based role cannot be.
You'll be hands-on when the situation calls for it, traveling to sites for complex escalations, new platform bring‑ups, and deployment support. But your primary leverage is organizational: building the technical standards, OEM relationships, and institutional knowledge that keeps Crusoe's GPU fleet reliable across every site.
What You'll DoCross‑Site Platform Operations & Escalation
Own Tier 2/3 hardware escalations across all Crusoe sites for issues that exceed local site capability, engaging directly with OEM and ODM engineering teams to drive resolution
Travel to sites as needed for complex platform issues, new hardware bring‑ups, and deployment support
Identify recurring failure patterns across sites and translate them into platform feedback, sparing strategy inputs, or OEM improvement requests
Root‑cause complex hardware issues — PCIe, BMC, thermal, fabric — and produce resolution documentation reusable across the Site Ops org
Hand off platform‑level findings to the appropriate internal engineering teams with clear, well‑documented escalation packages
OEM & ODM Technical Partnership
Develop and maintain deep technical relationships with Crusoe's primary hardware partners — currently Super Micro and HPE, with upcoming ODM’s as growing platforms — at the engineering and field escalation level
Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements that benefit the full fleet
Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint
Support vendor evaluations and new platform qualifications in partnership with Site Ops and engineering leadership
Platform Standards & Org Development
Own OEM platform technical knowledge at the Site Ops org level — escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites
Own the development and maintenance of platform‑specific SOPs, runbooks, and field troubleshooting procedures for the Site Ops org, ensuring site teams have current, actionable documentation across all active hardware platforms
Design and deliver technical training for Site…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).