×
Register Here to Apply for Jobs or Post Jobs. X

Senior Cloud Ops Engineer - SRE

Job in Salt Lake City, Salt Lake County, Utah, 84193, USA
Listing for: CaseWorthy, Inc.
Full Time position
Listed on 2026-07-19
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Azure, Systems Engineer
Salary/Wage Range or Industry Benchmark: 120000 - 180000 USD Yearly USD 120000.00 180000.00 YEAR
Job Description & How to Apply Below
Position: Senior Cloud Ops Engineer I - SRE

Job Summary

Senior Cloud Ops Engineer I – SRE is a site reliability engineer with a passion for implementing processes, tools, and methodologies that optimize the reliability, performance, and scalability of production systems across the software development lifecycle (coding, deployment, maintenance, and updates). As a member of the Cloud Ops team, you will contribute to its mission of supporting the larger engineering organization by scaling site reliability engineering (SRE) practices with a shared platform of modern, secure, next‑gen, self‑serve components and services.

The Senior Cloud Ops Engineer I – SRE will complement the team by bringing knowledge of operating systems administration, automation, coding and scripting, and reliability engineering for workloads running in the cloud, specifically Microsoft Azure. This role owns service level indicators (SLIs), service level objectives (SLOs), and error budgets for critical systems, and drives observability using Datadog and Azure Monitor.

Responsibilities
  • Maintain and improve the company's Azure cloud infrastructure with a focus on reliability, availability, and performance.
  • Define, track, and report on SLIs, SLOs, and error budgets for production services.
  • Provide on‑call, after‑hours support on a rotating schedule, including incident response and escalation.
  • Lead and participate in blameless post‑mortems and retrospectives of production infrastructure incidents, driving corrective and preventive action items to closure.
  • Perform hosting tasks as needed including deployments and patching.
  • Automate infrastructure and configuration management using Infrastructure as Code (IaC).
  • Maintain and execute new tenant site provisioning.
  • Plan and execute database migrations to accommodate capacity plans.
  • Build and maintain observability tooling — metrics, logs, traces, dashboards, and alerting — using Datadog and Azure Monitor to enhance security, observability, and monitoring of workloads (compute, data storage, application integration elements).
  • Perform capacity planning and load/performance analysis to anticipate and prevent reliability issues before they impact customers.
  • Contribute to the organization's security audits and risk assessments.
  • Assist with vulnerability scans / penetration tests for internal and client systems.
  • Assist with identifying, documenting, and socializing application risks and vulnerabilities.
  • Use Azure services such as Azure Virtual Machines, Azure Kubernetes Service (AKS), Azure Functions, Azure SQL Database / Azure SQL Managed Instance, Azure Cosmos DB, Azure Blob Storage, Azure API Management, and Azure DNS, to name a few.
  • Also work with Azure Dev Ops Pipelines / Git Hub Actions, Azure Resource Manager (ARM) templates / Bicep, Checkmarx, Tenable, Datadog, Azure Monitor, and Atlassian products.
  • Ability to travel nationwide, up to 10% annually.
  • Perform other duties as assigned.
Required

Skills & Qualifications
  • 3+ years of experience working with public cloud infrastructure, specifically Microsoft Azure.
  • Background in site reliability engineering (SRE), Dev Sec Ops , or software development.
  • Hands‑on experience with observability and monitoring platforms, specifically Datadog and Azure Monitor.
  • Working knowledge of SRE fundamentals — SLIs, SLOs, error budgets, incident response, and blameless post‑mortems.
  • Experience with deployment pipeline automation tools.
  • Experience with scripting languages (Python, Power Shell).
  • Basic knowledge of SQL.
  • B.S. in IT, Computer Science, or related field.
Preferred

Skills & Qualifications
  • Experience with Azure‑hosted environments.
  • Familiarity with Windows systems administration.
  • Experience with container orchestration (Azure Kubernetes Service / Kubernetes).
  • Microsoft Azure certifications (e.g., AZ‑104, AZ‑400, AZ‑500) are a plus.
  • Datadog certification is a plus.
  • Experience with AWS or GCP is a plus, given transferable public cloud skills.
  • M.S. in IT, Computer Science, or related field.
#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary