×
Register Here to Apply for Jobs or Post Jobs. X

Senior Software Engineer; Compute Architecture

Job in New York, New York County, New York, 10261, USA
Listing for: Chris Baily
Full Time position
Listed on 2026-07-22
Job specializations:
  • Software Development
    DevOps, Cloud Engineer - Software
Salary/Wage Range or Industry Benchmark: 120000 - 160000 USD Yearly USD 120000.00 160000.00 YEAR
Job Description & How to Apply Below
Position: Senior Software Engineer (Compute Architecture)
Location: New York

Requirements

  • 5+ years of experience building and operating infrastructure or backend systems
  • Bachelor’s or Master’s degree in Computer Science or a related field, or equivalent practical experience
  • Strong proficiency in Go for building production services and tools
  • Experience designing and building gRPC and REST APIs
  • Experience with Kubernetes and containerized workloads in production environments
  • Familiarity with observability tooling such as Prometheus and Grafana
  • (Desirable) Experience working with GPU-based systems
  • (Desirable) Experience with low-level hardware management such as BMCs or Redfish
  • (Desirable) Experience operating large-scale distributed systems or high-throughput infrastructure
  • (Desirable) Experience collaborating with or contributing to open-source projects (for example, Go, Redfish)
  • We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match
What the job involves
  • As a Senior Software Engineer within our Compute Architecture organization, you will help build the software control plane for hardware lifecycle management across large-scale GPU data centers
  • The METALDEV team builds Go-based distributed services that bring infrastructure online, monitor production hardware health, automate safe operational workflows, and give operators the observability and control needed to manage GPU servers and rack-scale systems with reliability and confidence
  • This is a software-first role at the intersection of distributed systems, production reliability, and hardware-aware automation, ideal for engineers who want their code to operate real-world infrastructure at massive scale
  • Design, build, and operate Go-based services that manage the lifecycle of large-scale GPU data center infrastructure
  • Build automation for data center bring-up, hardware discovery, health monitoring, remediation, and production operations
  • Develop reliable APIs, services, and workflows for managing BMCs, firmware state, server health, and rack-level infrastructure
  • Improve observability, alerting, and operational tooling so production issues can be detected, understood, and resolved quickly
  • Translate incidents and hardware failure modes into software improvements that make the platform more resilient
  • Partner with hardware-adjacent, infrastructure, operations, and software teams to design systems that work safely at fleet scale
#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary