×
Register Here to Apply for Jobs or Post Jobs. X

Senior Site Reliability Engineer- Health- (US Citizen

Job in Madison, Dane County, Wisconsin, 53786, USA
Listing for: Oracle
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below
Position: Senior Site Reliability Engineer- Oracle Health- (US Citizen)
** Job Description*
* Location / Work Authorization / Clearance

**-    Role is based in the United States.*
* **-    U.S. citizenship required due to security clearance requirements.*
* **-    No visa sponsorship available.*
* **-    Must be able to obtain and maintain the required security clearance.*
* We are seeking a Site Reliability Engineer to help build and operate reliable, scalable cloud-native platforms and services that support Oracle Health's next-generation healthcare technology initiatives. This role works across platform engineering, cloud infrastructure, distributed systems, integration services, healthcare applications, and operational tooling.

The ideal candidate is passionate about reliability, automation, observability, and performance. You will collaborate closely with software engineering, product, security, operations, and customer-facing teams to design resilient systems, improve service health, respond to incidents, and support modernization efforts across healthcare workflows, interoperability, data platforms, and AI-driven capabilities.

This role bridges software engineering and operations: designing infrastructure and services for reliability, developing automation, improving monitoring and alerting, supporting incident response and root cause analysis, and contributing to secure, scalable systems running on Oracle Cloud Infrastructure.

** Responsibilities*
* ** Responsibilities*
* Design, build, test, and operate reliable cloud infrastructure, platform capabilities, and services on Oracle Cloud Infrastructure and legacy deployment models.

Partner with software engineering teams to develop scalable, resilient services, APIs, integrations, and distributed systems.

Forecast capacity needs, analyze service trends, and take proactive steps to ensure systems can support current and future workloads.

Monitor service health, availability, latency, performance, and capacity using observability and reporting tools.

Define and maintain meaningful SLIs, SLOs, KPIs, dashboards, alerts, and runbooks for production services.

Improve service resilience through backup and restore validation, disaster recovery planning, secrets handling, patching, and least-privilege access practices.

Participate in incident response, troubleshooting, root cause analysis, postmortems, and follow-up remediation.

Develop automation, scripts, and tooling to support provisioning, deployment, monitoring, metrics collection, mitigation, and remediation.

Support safe release practices, including CI/CD, infrastructure automation, canary or blue-green deployments, rollback planning, and operational readiness reviews.

Investigate and debug issues across applications, infrastructure, services, and dependencies to help teams meet service level objectives.

Identify performance bottlenecks and reliability risks, then recommend and implement improvements.

Collaborate with product managers, architects, engineers, security, operations, and customer teams to deliver secure, customer-focused healthcare solutions.

Support modernization efforts involving cloud-native architectures, healthcare interoperability, large-scale healthcare data platforms, and AI-enabled capabilities.

Communicate service health, operational risks, capacity concerns, and the potential impact of infrastructure, feature, or tooling changes.

Contribute to documentation, runbooks, incident records, operational standards, and knowledge sharing.

Participate in on-call rotations and operational support for production services.

** Required Skills*
* Linux and networking:
Processes, file systems, systemd, DNS, TCP/IP, TLS, HTTP, load balancers, proxies, and basic database behavior.

Kubernetes operations:
Deployments, Services/Ingress, Config Maps/Secrets, RBAC, resource requests/limits, probes, autoscaling, persistent storage, Helm/Kustomize, container troubleshooting, and effective kubectl troubleshooting.

Cloud and infrastructure-as-code:
Oracle Cloud Infrastructure OCI or other cloud experience (AWS, GCP, Azure) , IAM, networks, compute, managed Kubernetes, Terraform, and configuration automation such as Ansible.

Delivery engineering:
Git, Git Hub, container images/registries,…
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary