×
Register Here to Apply for Jobs or Post Jobs. X

Senior Software Engineer

Job in Redmond, King County, Washington, 98073, USA
Listing for: Microsoft Corporation
Full Time position
Listed on 2026-09-05
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations
Job Description & How to Apply Below
** Overview*
* Microsoft Azure's Artificial Intelligence and High‑Performance Computing (AI/HPC) organization powers some of the world's largest cloud‑native supercomputers used for frontier AI training, scientific computing, and large‑scale distributed simulations. Our team builds and operates hyperscale GPU clusters that consistently place Azure among global leaders in the Top
500, MLPerf, and Graph
500 benchmarks. By joining us, you step into the engineering core responsible for ensuring these systems remain reliable, performant, and ready for the next wave of AI innovation.

At this supercomputing scale, reliability and operational excellence are engineering challenges of their own. As a Senior Supercomputing Operations Engineer, you will own day‑to‑day operations of GPU interconnect fabrics and treating them as a single, mission‑critical reliability domain that directly impacts GPU availability, training throughput, and customer SLAs. You will lead incident triage and mitigation, debug complex fabric‑layer failures, and correlate telemetry across nodes, switches, SM behavior, and GPU subsystems to identify true root causes.

Your work will focus on resolving real production incidents at scale, improving operational readiness, and preventing recurrence through better tooling, automation, and deep systems understanding.

You will build and use state‑of‑the‑art tools to detect issues proactively, close operational gaps, and improve observability across our fabrics. You will contribute to TSGs, operational playbooks, and escalation guides while partnering with internal engineering teams and industry leading manufacturers to drive meaningful fixes. The solutions you develop and the operational improvements you drive will uplift the reliability of Azure's largest supercomputing deployments and directly support the most compute‑intensive AI workloads running in the cloud.

Microsoft's mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

** Responsibilities*
* + Act as the DRI for supercomputing clusters and GPU compute and interconnect fabric operations, ensuring GPU availability, service reliability, and AI training stability.

+ Lead incident triage, mitigation, recovery, and root cause analysis for compute and fabric-related production issues across large-scale AI infrastructure.

+ Perform deep, cross-stack debugging spanning hardware provisioning, GPU interconnect fabric, PCIe subsystems, and GPU interactions to identify and resolve complex failures.

+ Drive operational excellence by identifying systemic failure patterns and developing technical guidance, troubleshooting procedures, playbooks, and escalation frameworks.

+ Design and leverage automation, telemetry, and diagnostic tooling to improve issue detection, observability, debuggability, and mean time to mitigation (MTTM).

** Qualifications*
* Required Qualifications:

+ Bachelor's Degree in Computer Science or related technical field AND 4+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python

+ OR equivalent experience.

Other Requirements:

+ Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings:

+ Microsoft Cloud Background Check:
This position will be required to pass the Microsoft Cloud Background Check upon hire/transfer and every two years thereafter.

Preferred Qualifications:

+ Master's Degree in Computer Science or related technical field AND 6+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python

+ OR Bachelor's Degree in Computer Science or related technical field AND 8+ years technical engineering experience with coding in…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary