Senior Site Reliability Engineer; C#, .NET
Listed on 2026-07-26
-
IT/Tech
SRE/Site Reliability, Systems Engineer
Senior Site Reliability Engineer
Remote | US
(EST Preferred)
The WorkAre you an experienced Site Reliability Engineer who thrives at the intersection of software engineering and production operations? Do you take pride in keeping mission‑critical customer systems reliable under real-world operational pressure? Are you looking for an opportunity to own production reliability for a modern hybrid infrastructure platform spanning cloud, colocation, and edge environments?
If so, we have an exceptional opportunity for you.
Climavision is seeking a Senior Site Reliability Engineertocontributetowardsreliability, operational excellence, and production resilience for our customer‑facing platform and weather data services. This role is focused on ensuring our systems consistently meet demanding customer SLAs, including a 99.5% availability commitment for radar‑derived data services. A central focus of this role isestablishingmulti‑replica and multi‑cluster high availability across our .NET services, including hands‑on refactoring of C# code to make services safe to run as multiple instances and across clusters.
This is a hands‑on engineering role for someone who is equally comfortable debugging production .NET services, troubleshooting Kubernetes clusters, leading incident response, and improving operational maturity across the organization. The successful candidate will combine strong software engineering experience in C# / .NET with deep production operationsexpertiseand a disciplined approach to reliability engineering.
Climavision operates a hybrid infrastructure footprint spanning Microsoft Azure, colocation data centers, and edge Kubernetes clusters, deployed alongside weather radar systems.
This role will drive production reliability across Azure, colocation, and edge environments.
35%
Production Reliability Engineering
30% Application Reliability & .NET Service Architecture
20% Kubernetes Platform Reliability/Operations
15% Observability,Automation, and Operational Excellence
Primary Responsibilities- Own production reliability for Climavision’scustomer‑facing platform and radar‑derived weather data services across Azure, colocation, and edge Kubernetes environments.
- Contribute to the definition and improvement of SLIs, SLOs, alerting standards, and operational metrics used to measure platform reliability.
- Support and coordinateproductionincident response efforts, including troubleshooting, mitigation, communication, and postmortem analysis.
- Diagnose and resolve complex production issues across application services, Kubernetes infrastructure, storage, and distributed systems.
- Drive multi‑replica and multi‑cluster high availability across
Climavision’s.NET services. This includes working directly in the C# codebase to refactor services that are not currently safe torun as multiple replicas, addressing in‑process state, sticky scheduling assumptions, non‑idempotent operations, race conditions, and other patterns that prevent safe horizontal scaling, so that services can be deployed with multiple replicas, across multiple clusters, for high availability. - Contribute tothe multi‑cluster high‑availability strategy across
Climavision’shybrid fleet, including active‑active and active‑passive failover behavior, traffic routing, data replication considerations, and graceful degradation when a cluster becomes unavailable. - Operate and improve
Climavision’sself‑managed Kubernetes platform spanning cloud‑hosted, colocation, and edge clusters with a focus on availability, resiliency,recovery and operational performance - Ensure Kubernetes platform lifecycle activities including upgrades, patching, cluster health, node management, and production change management, are executed inamanner thatpreservesservice availability and minimizes customer‑facing risk
- Improve reliability and operational maturity of production platform services, including observability, autoscaling, ingress, and distributed storage. Partner with the teams responsible for the underlying networking and security primitives rather than owning those areas directly.
- Design and validate
Kubernetes workloads for resiliency, scalability, and operational efficiency, including…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).