Engineering Manager, Cloud Monitoring Services Platform
Listed on 2026-09-28
-
Management
Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.
We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.
We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.
If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.
About the Role:Crusoe builds cloud infrastructure for AI workloads. Cloud Monitoring Services owns observability across Crusoe Cloud: metrics, logs, alerting, and the telemetry agent that runs on every node in the fleet. Our customers run large, demanding AI training and inference workloads, and they depend on us for a clear, trustworthy view of what their infrastructure is doing.
We are hiring an Engineering Manager to lead the Platform team: the time series and log storage systems, and the query layer that serves every dashboard, API call, and investigation on top of them. This is a first-line management role reporting to the Engineering Manager for Cloud Monitoring Services. You will take direct people management responsibility for a team of 4 to 6 engineers, growing, and own how telemetry is stored, retained, and read back at fleet scale.
This is a people-first leadership role with real delivery stakes. You will bring the technical depth to guide hard calls and earn your team's trust, but your success is measured through what your team accomplishes. Query latency, retention, and storage cost are live tradeoffs your team will be making continuously, and customers feel all three directly. The role sits alongside peer managers who own telemetry collection and the ingestion pipeline, and together you run a platform that has to work end to end.
This is a full-time position.
- Grow and develop your team. Manage 4 to 6 engineers directly: 1:1s, career growth, performance, and team health, with the team growing over time.
- Own storage and query. Be accountable for the time series and log storage systems and the query layer on top of them, including how they behave under load and how much they cost to run.
- Own delivery. Plan and sequence work across a roadmap that mixes customer-facing feature work with storage and query infrastructure, and make honest calls early when a plan is at risk.
- Set the technical direction for your area. Partner closely with the Staff engineers who own technical direction across the platform, so they can focus on engineering rather than absorbing delivery and handoff work alone. Ask the hard questions and make sound tradeoff calls alongside your engineers.
- Keep the operational bar high. Query is on the critical path when something is wrong in the fleet, so availability, query performance, and correctness of what comes back are core to the job, not afterthoughts.
- Own the team's oncall rotation and the operational health of the storage and query stack.
- Manage cost and scale together. Retention…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).