×
Register Here to Apply for Jobs or Post Jobs. X

Senior Hardware Development Engineer, Cloud AI​/ML Server Team

Job in Seattle, King County, Washington, 98127, USA
Listing for: Socket.dev
Full Time position
Listed on 2026-09-04
Job specializations:
  • Engineering
    Hardware Engineer, Systems Engineer, Test Engineer
Salary/Wage Range or Industry Benchmark: 159200 - 215300 USD Yearly USD 159200.00 215300.00 YEAR
Job Description & How to Apply Below

Final date to receive applications:
Sep 1, 2026

AWS operates the world's largest fleet of GPU-accelerated servers powering AI/ML training and inference at cloud scale. Our team defines the server architectures, drives the hardware designs, and owns the fleet quality for these platforms — from component selection through datacenter operations. If you want to shape the physical hardware that frontier models train on, this is the role.

We are seeking a Cloud Hardware Development Engineer to define server architectures based on workload demand, translate them into detailed component specifications, and drive validation from PCBA bring-up through rack integration. You will lead ODM design partners through development and production, triage hardware issues across manufacturing and datacenters, and own fleet quality metrics post-launch.

What You Will Do

You will define the hardware that runs the world's largest AI training workloads. Your designs span electrical, thermal, mechanical, power, and signal integrity across GPU-accelerated platforms. You will drive validation from first silicon through fleet-scale deployment, triage failures correlating across PCIe, power delivery, memory, and accelerator interconnects, and feed root cause findings back into design improvements. When a new server platform launches at a large scale, the architecture, component choices, and quality gates are yours.

Why

You Will Love It

The world's most advanced frontier models train on the hardware you design. You will see your architecture decisions scale to a large fleet of servers. The team is deeply technical and high-trust — you own platforms end to end from architecture definition through fleet operations.

The Ideal Candidate

You think across the full hardware stack — from silicon packaging and power delivery to rack-level thermal and mechanical design. You are as comfortable reviewing a schematic as you are analyzing fleet failure data. You drive quality through data, not assumption, and you hold design partners to the same standard you hold yourself. You mentor and develop junior engineers, contribute to hiring, and share your expertise to make the team stronger.

Key

job responsibilities Architecture & Design
  • - Define server architectures based on workload demand and customer requirements, translating them into detailed designs and component specifications that enable high-performance AI training and inference at scale
  • - Work with interdisciplinary teams of component, firmware, test, qualification, and integration engineers to deliver cohesive designs
  • - Drive design reviews with ODM/JDM partners covering schematic, layout, BOM, and manufacturing DFx (Design for Test, Design for Manufacturing)
Validation & Bring-up
  • - Define and execute validation strategies from PCBA bring-up through server and rack integration — covering power sequencing, signal integrity, thermal characterization, and accelerator interconnect performance
  • - Own hardware debug during EVT/DVT/PVT builds, correlating failures across PCIe, power rails, memory channels, and GPU subsystems
  • - Triage hardware issues at both ODM facilities and datacenters, conduct root cause analysis, and implement corrective actions
Fleet Quality & Continuous Improvement
  • - Own fleet quality metrics post-launch: server-level annualized failure rates and component-level failure modes
  • - Monitor operational telemetry to identify systemic issues and drive design or process changes for current and future platforms
  • - Partner with test and automation teams to improve manufacturing yield and reduce test dwell times
Cross-Team Collaboration
  • - Work with EC2 architecture teams to align on instance definitions, workload requirements, and platform trade-offs
  • - Drive ODM/JDM design partners through development milestones and production ramp
  • - Collaborate with firmware, software, and operations teams to ensure designs are debuggable, serviceable, and automation-ready

May require occasional (

Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary