Lead Principal Site Reliability Engineer - (work in Vienna VA location
Job in
Madison, Dane County, Wisconsin, 53786, USA
Listed on 2026-08-05
Listing for:
Oracle
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below
** Job Description*
* We are seeking a highly motivated Site Reliability Engineer (SRE) to support a strategic customer's cloud platform and mission-critical applications. The SRE will be responsible for ensuring high availability, operational excellence, automation, and continuous improvement of production environments.
The successful candidate will partner with application development, platform engineering, cloud infrastructure, networking, cybersecurity, and operations teams to build resilient, secure, and highly automated systems. This role emphasizes reliability engineering, observability, incident response, capacity planning, and proactive operational improvements.
Serves as a consultant and leads the design and architecture of infrastructure and service, ensuring alignment with reliability and functionality standards. Takes full ownership of forecasting of demands and responding to capacity needs. Owns collaborations with software development teams to develop reliable and scalable infrastructures. Oversees incident response and/or maintenance tasks. Provides strategic, future-oriented health and performance reporting. Contributes to strategies for automation and reviews the development and implementation of automation.
Leads the implementation of innovative tools and provides expertise in site reliability trends.
** Responsibilities*
* ** Key Responsibilities*
* + Maintain high availability, reliability, and performance of enterprise applications and cloud infrastructure.
+ Design and implement automation to reduce manual operational effort and improve deployment consistency.
+ Develop Infrastructure as Code (IaC) using Terraform and automation scripts using Python, Bash, or similar languages.
+ Build and maintain CI/CD pipelines to support automated deployments and release management.
+ Monitor applications and infrastructure using observability platforms, including metrics, logs, traces, and alerting.
+ Define and maintain Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
+ Participate in on-call rotations and respond to production incidents with urgency and professionalism.
+ Conduct root cause analysis (RCA) and implement corrective and preventive actions to eliminate recurring issues.
+ Perform capacity planning, performance tuning, and scalability assessments.
+ Support Kubernetes clusters, containerized workloads, and cloud-native applications.
+ Collaborate with development teams to improve application resiliency, fault tolerance, and operational readiness.
+ Partner with security teams to ensure systems comply with enterprise security and regulatory requirements.
+ Create operational runbooks, documentation, and standard operating procedures.
+ Continuously improve platform reliability through automation, monitoring, and operational best practices.
+ Takes full ownership of forecasting infrastructure demands and strategically responds to capacity needs, ensuring systems have sufficient resources to handle current and future workloads, anticipating resource gaps.
+ Seeks opportunities for prototyping and encourages the implementation of prototyping initiatives (e.g., testing new applications or infrastructures, assisting in onboarding), driving new approaches.
+ Identifies and projects resource gaps using cost calculators, identifying alternative solutions to lower costs, when necessary.
+ Provides guidance to others when monitoring services, maintains up-to-date knowledge of their performance, and meticulously documents their condition.
+ Oversees root cause analyses for incidents and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery), ensuring efficient execution and preventing incident reoccurrence.
+ Provides strategic, future-oriented health and performance reporting, and anticipates actions needed based on trends in data.
+ Provides expert-level release notes, communication, and/or guidance on the scale, capacity, security, performance attributes, and requirements of services and technology to customers, cross-functional teams, leadership, and external stakeholders.
+ Provides leadership in the on-call shifts.
+ Leads the resolution of multifaceted technical issues spanning multiple services, functions, and customers, collaborating with cross-functional teams and leveraging advanced investigation and debugging techniques to ensure the achievement of SLOs (service level objectives).
+ Provides input on strategic initiatives to address opportunities to improve performance bottlenecks and deployments, maximizing resource utilization, cost efficiency, speed, and scalability across their line of business.
+ Provides expertise in site reliability trends, guiding the creation and sharing of insights and best practices to shape the future of building, testing, deploying, and running services.
** Required Qualifications*
* + Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×