×
Register Here to Apply for Jobs or Post Jobs. X

Senior System Reliability Engineer

Job in Redmond, King County, Washington, 98053, USA
Listing for: Microsoft Corporation
Full Time position
Listed on 2026-07-19
Job specializations:
  • Engineering
    Systems Engineer, Electrical Engineering
Job Description & How to Apply Below
Overview

Microsoft Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft's expanding Cloud Infrastructure and responsible for powering Microsoft's "Intelligent Cloud" mission. SCHIE delivers the core infrastructure and foundational technologies for Microsoft's over 200 online businesses including Bing, MSN, Office 365, Xbox Live, Teams, One Drive, and the Microsoft Azure platform globally with our server and data center infrastructure, security and compliance, operations, globalization, and manageability solutions.

Our focus is on smart growth, high efficiency, and delivering a trusted experience to customers and partners worldwide and we are looking for passionate engineers to help achieve that mission.

As Microsoft's cloud business continues to grow the ability to deploy new offerings and hardware infrastructure on time, in high volume with high quality and lowest cost is of paramount importance. To achieve this goal, the Hardware, Infrastructure Management, and Fundamentals Engineering (HIFE) team is instrumental in defining and delivering operational measures of success for hardware manufacturing, improving the planning process, quality, delivery, scale and sustainability related to Microsoft cloud hardware.

We are looking for engineers with a dedicated passion for customer focused solutions, insight and industry knowledge to envision and implement future technical solutions that will manage and optimize the cloud infrastructure.

We are looking for a Senior System Reliability Engineer to join the team.

#SCHIE

Responsibilities

* Ability to drive both Design for Reliability (DfR) and Reliability, Availability & Maintainability (RAM) processes with consistency and rigor for both current and next generation Cloud & AI hardware and infrastructure solutions.

* Work across functionally across different disciplines (Hardware, Firmware, Architecture, Safety, Serviceability etc.) to influence business and operational decisions.

* Works with key stakeholders to set appropriate reliability and availability targets and allocates the targets to lower-level sub-systems and components.

* Develops appropriate tests at system, sub-system and/or/component as needed to demonstrate the budgeted reliability targets.

* Carries out effective Design Failure Modes & Effects Analysis (DFMEA) to identify and mitigate critical risks via design, operational and diagnostic improvements.

* Utilizes different modeling methods such as Reliability Block Diagram (RBD), Markov, Discreet Event Simulations (DES) and Fault Tree Analysis (FTA) to quantify risks and/or support business decisions.

* Develops Prognostics & Health Management (PHM) models for Remaining Useful Life (RUL) prediction based on telemetry data.

Qualifications

Required Qualifications

* Doctorate Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 2+ years technical engineering experience OR Master's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 4+ years technical engineering experience OR Bachelor's Degree in Mechanical Engineering, Materials Engineering, Reliability Engineering, Electrical Engineering, or related field AND 5+ years technical engineering experience OR 12+ years relevant technical engineering experience.

Other requirements

Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include but are not limited to the following specialized security screenings:
Microsoft Cloud Background Check:
This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.

Preferred Qualifications

* M.S. in Electrical or Electronic Engineering, Reliability Engineering, or equivalent discipline

* 5+ years of experience performing Reliability Engineering on Complex Systems

* Background in Applied Reliability Engineering statistics for repairable and non-repairable systems.

* Experienced in using software tools such as Reliasoft, JMP and python scripts for…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary