Senior Firmware Engineer
Job in
Redmond, King County, Washington, 98053, USA
Listed on 2026-07-20
Listing for:
Microsoft Corporation
Full Time
position Listed on 2026-07-20
Job specializations:
-
Software Development
Embedded Systems/ Firmware/ IoT, Cloud Engineer - Software, DevOps, Software Engineer
Job Description & How to Apply Below
Are you passionate about building highly reliable cloud infrastructure on a planetary scale? The Azure Hardware Diagnostics team is responsible for delivering firmware-driven diagnostic and validation solutions that ensure the health, reliability, and availability of Azure datacenter hardware.
We are seeking a Senior Firmware Engineer with deep expertise in platform firmware, hardware diagnostics, and failure analysis to drive end-to-end diagnostic capabilities across Azure server platforms. In this role, you will be a technical leader responsible for architecting, developing, and deploying firmware-based diagnostic solutions that enable rapid fault isolation, root-cause analysis, and predictive health monitoring for compute, memory, storage, networking, and accelerator subsystems.
You will work across the hardware, firmware, operating system, and cloud-stack boundaries, partnering closely with silicon vendors, platform engineering teams, datacenter operations, reliability engineering, and Azure service teams to improve hardware quality and reduce customer impact from infrastructure failures. Your work will directly influence the reliability and availability of millions of servers powering Microsoft's cloud services.
Responsibilities
* Own firmware-based diagnostic readiness across Azure server platforms from architecture and bring-up through production support.
* Design diagnostic frameworks across silicon telemetry, hardware interfaces, platform firmware, operating systems, and cloud health systems.
* Lead platform bring-up, validation, and root-cause investigations for complex hardware, firmware, and system software issues.
* Develop diagnostic telemetry, health monitoring, fault-detection, automation, and triage tooling to improve fleet reliability.
* Collaborate with silicon vendors, ODM/OEM partners, and internal platform teams to define diagnostic requirements and improve serviceability.
* Mentor engineers and drive engineering excellence through design reviews, code reviews, metrics, validation coverage, and secure firmware practices.
Qualifications
Required Qualifications
* Doctorate in Electrical Engineering, Computer Engineering, Computer Science, or related field AND 1+ year(s) technical engineering experience OR Master's Degree in Electrical Engineering, Computer Engineering, Computer Science, or related field AND 4+ years technical engineering experience OR Bachelor's Degree in Electrical Engineering, Computer Engineering, Computer Science, or related field AND 5+ years technical engineering experience
* OR equivalent experience
Other Requirements
* Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:
Microsoft Cloud Background Check:
This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.
Additional qualifications
* Experience with large-scale cloud infrastructure, hyperscale datacenters, Azure infrastructure, or fleet health management systems.
* 6+ years developing platform firmware, embedded software, or low-level systems software for server, embedded, or datacenter platforms.
* Strong proficiency in C/C++ with experience building, debugging, and shipping firmware solutions through full development cycles.
* Experience developing or debugging firmware-based diagnostic and validation solutions for server-class hardware.
* Ability to independently root-cause cross-stack issues involving hardware, firmware, operating systems, telemetry, and infrastructure.
* Strong collaboration and technical leadership skills, including driving discussions across multiple engineering organizations and external partners.
* Knowledge of memory diagnostics, PCIe diagnostics, storage reliability, accelerator/GPU health monitoring, or platform serviceability.
* Experience with firmware telemetry pipelines, cloud-scale diagnostic data analytics, reliability analysis, or predictive failure detection.
* Experience developing automation frameworks for validation, failure triage,…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×