More jobs:
Senior Manager, Site Reliability Engineering
Job in
Madison, Dane County, Wisconsin, 53786, USA
Listed on 2026-08-05
Listing for:
Oracle
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, Cybersecurity, SRE/Site Reliability
Job Description & How to Apply Below
* Supports team members in designing and architecting infrastructure and service and shares guidance on practices for reliability and functionality. Provides direction to ensure accurate forecasting and ensure systems have adequate resources, identifying resource gaps. Maintains a collaborative relationship with the software development team to create reliable, scalable infrastructures. Monitors data collection and ensures team members optimize operations and infrastructure reliability. Aids in incident response activities to ensure service reliability.
Monitors health and performance reports. Implements standards for identifying and recommending automation. Ensures team members communicate information and articulate the impact of changes. Serves as a senior management point and shares expectations for documentation. Sets expectations for experimenting with new technology, executing improvements, building site reliability knowledge, and providing clear data.
** Responsibilities*
* ** Key Responsibilities*
* ** Capacity Ingestion and Management:*
*
- Supports
team members designing and architecting infrastructure and/or service, sharing
guidance on practices and terms for reliability and functionality.
- Supervises
team members and provides direction to ensure accurate forecasting of demands
for infrastructure and response to capacity needs, ensuring systems have
sufficient resources to handle current and future workloads and identifying
resource gaps.
- Maintains
a collaborative relationship with the software development team to develop
infrastructures, ensuring features are reliable and scalable according to
deployment requirements.
- Implements
expectations for identifying opportunities for prototyping and manages
prototyping initiatives (e.g., testing new applications or infrastructures,
assisting in onboarding) to explore novel approaches.
** Incident and Service Lifecycle Management:*
*
- Monitors
data collection, triage, technical analysis, and redirection, ensuring team
members maintain and optimize operations and infrastructure reliability.
- Provides
support to team members monitoring services, ensuring they maintain up-to-date
knowledge of performance and document their condition.
- Leverages
advanced knowledge to aid team members in performing incident response, root
cause analyses, and/or maintenance on assigned services (e.g., software
installs, version upgrades, security updates, backup and recovery).
- Monitors
comprehensive health and performance reporting and ensures team members take
appropriate actions based on trends in data.
- Ensures
team members adhere to procedures when performing provisioning to support
infrastructure, applications, and services.
- Encourages
team members to experiment with new approaches for and perform decommissioning
(e.g., shutting down servers, removing data from databases) to remove objects
that are no longer needed.
** Automation:*
*
- Implements
standards for identifying and recommending opportunities for automation and
assesses potential benefits to enhance operational efficiency.
- Takes a
proactive role in reviewing and offering feedback on design, automation tools,
or scripts, acting as a leader during implementation.
- Shares
strategies for conducting testing on automations to ensure they perform tasks
correctly and produce expected results.
** Technical Communication and Guidance:*
*
- Reviews
and provides feedback on release notes and ensures team members communicate
comprehensive information about the scale, capacity, security, performance
attributes, and requirements of services and technology with customers and
immediate and related teams.
- Proactively
anticipates and articulates the potential impact of infrastructure, feature,
and tool changes, considering their impact across team operations.
- Serves
as a resource to team members on what information to communicate and how to
communicate.
** Troubleshooting and Resolution:*
* - Serves as a senior
management escalation point for incidents and complex issues arising within
Oracle services.
- Monitors the resolution of
technical issues spanning multiple services, ensuring effective investigation
and debugging…
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×