Site Reliability Engineer — AI Accelerator Infrastructure
Listed on 2026-07-16
-
IT/Tech
Hardware Engineer, Systems Engineer, IT Support
At d-Matrix
, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.
We value humility and believe in direct communication. Our team is inclusive
, and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground?
Together
, we can help shape the endless possibilities of AI.
d-Matrix designs purpose-built AI inference silicon. Our unified infrastructure team, SRE engineers, Dev Ops engineers, and data center & lab technicians own the physical and virtual layer that every engineering team depends on. As a DC & lab technician, you are the hands and feet of that team: racking servers, running cables, executing hardware bring-ups, and keeping lab environments in the precise, audit-ready state that high-velocity silicon and software development demands.
Aboutthe Role
This is a one-year contract with potential for a full-time conversion. This is an ownership role, not a ticket executor role. You operate independently, document what you build, and elevate with precision when something needs engineering attention.
What You Will Do Physical Infrastructure & Hardware Bring-Up- Rack, stack, cable, and decommission servers, PDUs, network gear, and storage in on-premises labs and colocation facilities.
- Execute hardware bring-up for d-Matrix AI-accelerated systems and validation test benches, including BIOS configuration, firmware validation, and OS installation (Linux primary).
- Replace and upgrade components: PCIe cards, NICs, storage, memory, and accelerator hardware across multiple server generations.
- Maintain lab spaces in organized, ESD-compliant, and audit-ready condition at all times.
- Own accurate asset tracking and inventory records using DCIM tools (Net Box, Jira, or equivalent); every piece of hardware is accounted for with the current configuration state.
- Coordinate equipment moves between lab, staging, and colocation; manage shipping, receiving, and RMA workflows.
- Track hardware lifecycle status: warranty, EOL, refresh schedules, and spare parts inventory.
- Diagnose and resolve hardware failures, power and thermal issues, and network connectivity problems – escalating to SRE engineers with clear documentation and reproduction steps.
- Serve as the first physical responder for infrastructure incidents requiring hands-on intervention in lab or colo environments.
- Document all troubleshooting actions in the team’s ticketing and knowledge base systems, contributing to runbooks that reduce repeat escalations.
- Write and maintain SOPs, rack diagrams, and cabling guides clear enough for a new technician to execute without shadowing.
- Communicate issues and status across hardware, software, SRE, and silicon validation teams professionally and with precision.
- Support parallel work streams as silicon development programs evolve; operate independently when SREs are heads-down on engineering work.
Minimum Qualifications
- 3+ years in data center operations, hardware validation, or systems technician roles, hands-on with real servers.
- Proven rack-and-stack experience: physical server installation, PDU and patch panel cabling, rack power planning, and cable management to professional standards.
- Hardware bring-up experience: BIOS/UEFI configuration, firmware updates, component replacement, and OS installation across Linux distributions.
- Linux command-line proficiency: enough to run diagnostics, inspect logs, and execute runbooks independently.
- Asset management discipline: experience with DCIM or inventory tools and a track record of accurate, up-to-date records.
- Strong written communication: you document what you do, escalates with context, and writes SOPs others can follow.
- Comfortable operating independently in a fast-moving startup…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).