×
Register Here to Apply for Jobs or Post Jobs. X

Principal Site Reliability Engineer

Job in Madison, Dane County, Wisconsin, 53786, USA
Listing for: Oracle
Full Time position
Listed on 2026-08-05
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer
Job Description & How to Apply Below
** Job Description*
* As a Principal Site Reliability Engineer (IC4), you will be responsible for designing, building, and operating highly available, scalable, secure, and resilient cloud services. You will combine software engineering with infrastructure expertise to improve service reliability, operational efficiency, and developer productivity across large-scale distributed systems.

You will lead complex reliability initiatives, drive automation-first operational practices, and develop software solutions that eliminate manual toil. You will partner closely with software engineering, cloud infrastructure, security, and product teams to architect resilient platforms that meet aggressive availability, scalability, and performance objectives.

Success in this role requires deep expertise in distributed systems, cloud infrastructure, coding, automation, observability, incident management, and operational excellence. You will leverage modern AI technologies, machine learning, and intelligent automation to streamline operations, accelerate incident response, improve troubleshooting, and enable autonomous system management.

You are expected to be a technical leader who influences architecture, establishes engineering best practices, mentors other engineers, and drives continuous improvements across multiple services and organizations.

** Responsibilities*
* ** Reliability Engineering & Service Ownership*
* + Design, build, and operate highly available, scalable, and fault-tolerant cloud services that meet defined Service Level Objectives (SLOs) and Service Level Agreements (SLAs).

+ Lead architecture reviews to improve resiliency, scalability, observability, and operational readiness.

+ Forecast infrastructure growth, capacity requirements, and resource utilization while proactively mitigating operational risks.

+ Continuously improve platform reliability through engineering solutions rather than manual operational processes.

+ Define reliability standards, operational best practices, and service readiness criteria across multiple engineering teams.

** Software Engineering, Automation & AI*
* + Design and develop production-quality software, automation frameworks, and internal platforms using  
** Python, Java, Go, or similar programming languages** .

+ Build scalable automation to eliminate repetitive operational work, reduce manual intervention, and improve engineering productivity.

+ Develop APIs, microservices, and tooling that simplify infrastructure management, deployment, monitoring, and operational workflows.

+ Leverage  
** Generative AI, Large Language Models (LLMs), AI agents, and intelligent automation
** to:

+ Automate routine operational tasks and runbooks.

+ Accelerate incident triage and root cause analysis.

+ Improve log analysis and anomaly detection.

+ Generate operational insights and recommendations.

+ Automate knowledge management and operational documentation.

+ Enhance developer productivity and self-service capabilities.

+ Identify opportunities to incorporate AI-driven operational intelligence into existing systems to improve efficiency, reliability, and scalability.

** Infrastructure Engineering*
* + Design and optimize cloud infrastructure supporting distributed services across multiple regions and availability domains.

+ Improve system resiliency through redundancy, automation, and infrastructure-as-code.

+ Build and maintain deployment pipelines, provisioning frameworks, and configuration management solutions.

+ Drive infrastructure standardization and platform modernization initiatives.

** Observability & Operational Excellence*
* + Design comprehensive monitoring, logging, tracing, and alerting strategies.

+ Build meaningful dashboards, health reporting, and service performance metrics.

+ Improve alert quality, reduce operational noise, and enhance system visibility.

+ Define and measure Service Level Indicators (SLIs), SLOs, and error budgets.

+ Continuously optimize operational processes using data-driven insights.

** Incident Management & Reliability*
* + Lead critical production incident response and act as a senior escalation point during major service events.

+ Drive root cause analysis, corrective actions, and post-incident reviews with a focus on long-term engineering improvements.

+ Develop automated remediation solutions to reduce Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).

+ Continuously improve operational readiness through disaster recovery testing, game days, and failure injection exercises.

** Performance & Scalability*
* + Identify performance bottlenecks across applications, infrastructure, storage, databases, and networking.

+ Drive optimization initiatives to improve system throughput, latency, efficiency, and cost.

+ Perform capacity planning and predictive scaling using historical trends and telemetry data.

+ Optimize resource utilization while maintaining service reliability and customer experience.

** Technical Leadership*
* + Serve as the technical leader for reliability engineering…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary