HPC Data Center Developer
Listed on 2026-07-18
-
IT/Tech
IT Infrastructure, Network Engineer
HPC Data Center Production Engineer
Chicago, IL or New York, NY – On-site 5 days/week.
Jump Trading Group is a global organization of engineers who architect, build, and maintain world‑class trading infrastructure. We empower exceptional talents in Mathematics, Physics, and Computer Science to push scientific boundaries and apply cutting‑edge research to global financial markets. Our culture demands fearlessness, creativity, intellectual honesty, and relentless competitiveness.
We are looking for an HPC Data Center Production Engineer to build and own the automation and tooling that powers Jump’s HPC data‑center operations. This development‑heavy role focuses on automating the onboarding and lifecycle management of data‑center hardware—servers, switches, rack PDUs, CDUs, and environmental sensors—and building tools for capacity planning, outage simulation, monitoring, and metrics integration.
Responsibilities- Hardware Onboarding Automation – Design, develop, and maintain automation to onboard new hardware devices—including servers, network switches, rack PDUs, CDUs, and environmental sensors—taking them from racked and cabled to discovery, configuration, validation, and production‑ready state with minimal manual intervention.
- Data Center Tooling Development – Build tools for power and cooling capacity planning, outage simulation, and day‑to‑day operational support such as hardware lifecycle tracking, inventory management, change management, and diagnostics.
- Monitoring & Metrics Integration – Pull telemetry from all infrastructure components into centralized observability platforms, integrate provider metrics feeds, and implement the monitoring and alerting strategy.
- Cross‑Team Collaboration – Work closely with HPC Planning, Engineering, and Operations leads to translate tooling and monitoring needs into production‑ready systems, and partner with HPC Engineering on integration points with compute, storage, and network provisioning.
- Systems Maintenance & Reliability – Own reliability and lifecycle of all developed systems, monitor for failures, respond to incidents, iterate based on feedback, maintain comprehensive documentation, and participate in large maintenance operations including evenings and weekends.
- AI‑Driven Development – Use AI tools daily for code, debugging, documentation, and accelerate development velocity. Identify opportunities to apply AI to data‑center operations such as anomaly detection and predictive planning.
- Perform additional duties as assigned or needed.
- 5+ years of professional experience in production engineering, infrastructure automation, or site reliability engineering, preferably in HPC or large‑scale data‑center environments.
- Proven track record of building and shipping production automation and tooling that is maintained and reliable.
- Experience automating hardware provisioning and lifecycle management for servers, network devices, and power/cooling infrastructure.
- Strong understanding of data‑center infrastructure: power distribution, cooling systems (air and liquid), environmental monitoring, and structured cabling.
- Experience integrating with hardware management interfaces (IPMI, BMC, Redfish, SNMP, vendor APIs) for discovery, configuration, and telemetry collection.
- Result‑driven, high‑energy professional able to work under pressure with tight deadlines.
- Excellent written and verbal communication skills.
- Reliable and predictable availability, including ability to work evenings and weekends as required.
- Bachelor’s degree preferred.
- High proficiency in Golang and at least one additional language (e.g., Python).
- Strong Linux systems knowledge, including system administration, networking, storage, process management, log analysis, and troubleshooting at the OS level.
- Experience with Grafana dashboards, Prometheus, InfluxDB, or similar observability platforms and building custom integrations or exporters.
- Experience with configuration‑management and infrastructure‑as‑code tools such as Salt Stack, Ansible, Terraform.
- Solid networking knowledge: L2/L3 protocols, VLANs, BGP, SNMP, and switch/router configuration (Arista, Cisco).
- Experience consuming…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).