×
Register Here to Apply for Jobs or Post Jobs. X

Site Reliability Engineer - CTJ - POLY with Security Clearance

Job in Redmond, King County, Washington, 98052, USA
Listing for: Microsoft Corporation
Full Time position
Listed on 2026-08-01
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, Systems Engineer, AI Engineer (Applied/Software), Machine Learning/ ML Engineer
Job Description & How to Apply Below
Overview The Silver Edge team brings the power of Azure to the edge for our customers, tackling some of the most complex and mission-critical challenges in cloud and edge computing. Our mission is to provide stellar customer service so that their mission can succeed. We support the new Azure Local product that brings cloud computing to local hardware. We're looking for a new member of our team that relishes solving complex, ambiguous problems at scale and is passionate about building resilient systems that matter.

As a Site Reliability Engineer on the Silver Edge team, you will work on building out and ensuring the dependability of Azure Local services in 3 different sovereign clouds. You will be required to solve tough technical problems, and thrive in dynamic, sometimes chaotic environments. In this role, you will accelerate your career, deepen your expertise in sovereign cloud solutions and help implement the future of Azure edge solutions.

We offer flexible work arrangements, including partial remote options, to support your best work. Microsoft's mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others, and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.

Responsibilities Primary Responsibilities
• Support customer deployments and use of Azure Local and Azure Local disconnected operations.
• Maintain Azure Service reliability including deployment, availability, security, performance and customer satisfaction for sovereign environments. Contributions to Development and Design
* Leverages technical expertise in cloud technologies and specific products, as well as objective insights drawn from analyses of production telemetry data to suggest changes or add-ons to product features or the automation to improve the availability, security, quality, observability, reliability, efficiency, observability, and performance of product components or features supported by their team.

* Engages with product engineering teams by participating code/design reviews, regular meetings, on-call rotations and incident responses throughout product development and operations cycles. Utilizes technical knowledge of systems/platforms and insights drawn from product engineering teams, security best practices, artificial intelligence (AI)/machine learning (ML), and telemetry analyses to suggest potential improvements in code base and designs across components and features of one or more products.

Driving Operational Excellence
* Leverages technical expertise and telemetry analysis alongside advanced artificial intelligence (AI) and machine learning (ML) algorithms across a range of components and/or features to identify patterns and opportunities to implement configuration and data changes for one or more platforms, systems, or products in production using code, tooling, and automation.

* Independently writes code or scripts that automate the performance of scalable operations processes (e.g., monitoring, alerting, deploying products and updates) across components and features of products operating at scale.

* Shares insights and best practices via documented artifacts that can be applied to improve development and operations of system, platform, or product components and features by participating in code/design reviews, incident drills and debriefs, and regular meetings, as well as interactions with more experienced SREs and members of product engineering teams.

* Develops alerts and instrumentation across components and features to monitor product capacity, related security risk, and resource demands and analyze telemetry data using existing capacity planning models. Draws insights from analyses of capacity and resource data to optimize component and feature code to manage resources and capacity across limited range of use conditions and system parameters.

* Independently uses existing tools and/or models to troubleshoot problems or flaws affecting the availability, security, reliability, performance, and/or efficiency of components and features, leveraging the artificial intelligence (AI) and machine learning (ML) capabilities. Proposes solutions that will resolve and prevent recurring issues and brings them to the attention of their Site Reliability Engineering (SRE) and/or product engineering teams.

* Utilizes insights from performance and resource monitoring tools to identify whether there is a need to optimize the efficiency of component and feature code, or if changes to compute resources are required. Models the predicted effect of changes to code and/or compute resources across components or features to document the efficacy of proposed solutions. Proposes changes and drives implementation of solutions to identified performance and resource challenges.

* Identifies…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary