×
Register Here to Apply for Jobs or Post Jobs. X
More jobs:

Senior Staff Software Engineer, DC Infrastructure

Job in Northern, Floyd County, Kentucky, USA
Listing for: CV in
Full Time position
Listed on 2026-08-21
Job specializations:
  • Software Development
    DevOps
Salary/Wage Range or Industry Benchmark: 170000 - 220000 USD Yearly USD 170000.00 220000.00 YEAR
Job Description & How to Apply Below
Location: Northern

Senior Staff Software Engineer, DC Infrastructure at crusoe.

About the role Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We are seeking a highly skilled and motivated Software Engineer to join Crusoe’s Data Center Infrastructure Engineering team. This role is focused on the development of software for the management of a fleet of GPU servers as well as the data centers that house those systems. The ideal new team member will be a hands‑on problem solver who is comfortable working independently.

The new team member will play a critical role in maintaining the health and scalability of Crusoe’s rapidly growing GPU fleet. You will own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success. The team owns deployment maintenance, observability, critical environment management, and automation. You will develop automation and AI agents for executing component‑level diagnosis and remediation for failed or degraded hardware.

Key

facts

Location:

San Francisco, CA - US Engagement:

What you’ll do
  • Developing and implementing deep‑level diagnostics and troubleshooting of hardware faults within GPU racks and high‑density compute systems.
  • Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
  • Developing automation and AI agents for executing component‑level diagnosis and remediation for failed or degraded hardware.
  • In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.
  • Developing tooling for post‑repair validation and testing tools such as burn‑in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.
  • Owning the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
  • Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems.
  • Writing and maintaining scalable, observable, and resilient software that integrates with existing cloud and on‑premise infrastructure.
  • Collaborating closely with cross‑functional teams to define, build, and deliver features that meet strict reliability and performance targets.
  • Contributing to on‑call rotations to support critical production issues and driving rapid resolution.
  • Implementing observability pipelines, metrics, and alerting to provide deep insight into system health and performance.
  • Refactoring legacy scripts and workflows into robust, maintainable services that improve reliability and developer velocity.
  • Participating in design reviews and ensuring best practices for security, scalability, and maintainability are followed.
  • Mentoring junior engineers through code reviews, pair programming, and knowledge sharing sessions.
Requirements
  • Software engineering experience.
  • The ability to identify a problem, rapidly develop a scalable solution and ship it.
  • Ability to lean in and assist team members working on critical or complex technical initiatives.
  • Ability to set the technical direction for a specific project and execute.
  • Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.).
  • Strength in at least one programming language - Go, Python, Java, Rust.
  • Strong analytical and problem‑solving skills.
  • Excellent communication and collaboration skills.
  • Ability to work independently and within a team.
Nice to have
  • Experience with Temporal and Kubernetes.
  • Experience working directly with hardware vendors.
  • Background in large‑scale GPU fleet operations or hyperscale data center environments.
Practical notes

This role may require occasional travel to client sites or Crusoe offices. Candidates must be authorized to work in the United…

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary