×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer — AI Accelerator Infrastructure

Job in Santa Clara, Santa Clara County, California, 95053, USA
Listing for: d-Matrix inc.
Full Time, Contract position
Listed on 2026-07-16
Job specializations:
  • IT/Tech
    Hardware Engineer, Systems Engineer, IT Support
Salary/Wage Range or Industry Benchmark: 65000 - 90000 USD Yearly USD 65000.00 90000.00 YEAR
Job Description & How to Apply Below
Position: Contract Site Reliability Engineer — AI Accelerator Infrastructure

At d-Matrix
, we are focused on unleashing the potential of generative AI to power the transformation of technology. We are at the forefront of software and hardware innovation, pushing the boundaries of what is possible. Our culture is one of respect and collaboration.

We value humility and believe in direct communication. Our team is inclusive
, and our differing perspectives allow for better solutions. We are seeking individuals passionate about tackling challenges and are driven by execution. Ready to come find your playground?
Together
, we can help shape the endless possibilities of AI.

Data Center & Lab Technician - AI Accelerator Infrastructure - Contract

d-Matrix designs purpose-built AI inference silicon. Our unified infrastructure team, SRE engineers, Dev Ops engineers, and data center & lab technicians own the physical and virtual layer that every engineering team depends on. As a DC & lab technician, you are the hands and feet of that team: racking servers, running cables, executing hardware bring-ups, and keeping lab environments in the precise, audit-ready state that high-velocity silicon and software development demands.

About

the Role

This is a one-year contract with potential for a full-time conversion. This is an ownership role, not a ticket executor role. You operate independently, document what you build, and elevate with precision when something needs engineering attention.

What You Will Do Physical Infrastructure & Hardware Bring-Up
  • Rack, stack, cable, and decommission servers, PDUs, network gear, and storage in on-premises labs and colocation facilities.
  • Execute hardware bring-up for d-Matrix AI-accelerated systems and validation test benches, including BIOS configuration, firmware validation, and OS installation (Linux primary).
  • Replace and upgrade components: PCIe cards, NICs, storage, memory, and accelerator hardware across multiple server generations.
  • Maintain lab spaces in organized, ESD-compliant, and audit-ready condition at all times.
Asset Management & Inventory
  • Own accurate asset tracking and inventory records using DCIM tools (Net Box, Jira, or equivalent); every piece of hardware is accounted for with the current configuration state.
  • Coordinate equipment moves between lab, staging, and colocation; manage shipping, receiving, and RMA workflows.
  • Track hardware lifecycle status: warranty, EOL, refresh schedules, and spare parts inventory.
Troubleshooting & Incident Support
  • Diagnose and resolve hardware failures, power and thermal issues, and network connectivity problems – escalating to SRE engineers with clear documentation and reproduction steps.
  • Serve as the first physical responder for infrastructure incidents requiring hands-on intervention in lab or colo environments.
  • Document all troubleshooting actions in the team’s ticketing and knowledge base systems, contributing to runbooks that reduce repeat escalations.
Documentation & Collaboration
  • Write and maintain SOPs, rack diagrams, and cabling guides clear enough for a new technician to execute without shadowing.
  • Communicate issues and status across hardware, software, SRE, and silicon validation teams professionally and with precision.
  • Support parallel work streams as silicon development programs evolve; operate independently when SREs are heads-down on engineering work.
What You Will Bring

Minimum Qualifications
  • 3+ years in data center operations, hardware validation, or systems technician roles, hands-on with real servers.
  • Proven rack-and-stack experience: physical server installation, PDU and patch panel cabling, rack power planning, and cable management to professional standards.
  • Hardware bring-up experience: BIOS/UEFI configuration, firmware updates, component replacement, and OS installation across Linux distributions.
  • Linux command-line proficiency: enough to run diagnostics, inspect logs, and execute runbooks independently.
  • Asset management discipline: experience with DCIM or inventory tools and a track record of accurate, up-to-date records.
  • Strong written communication: you document what you do, escalates with context, and writes SOPs others can follow.
  • Comfortable operating independently in a fast-moving startup…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary