Director, Site Reliability Engineering
Job in
Sacramento, Sacramento County, California, 94278, USA
Listed on 2026-07-03
Listing for:
Oracle
Full Time
position Listed on 2026-07-03
Job specializations:
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Project Manager
Job Description & How to Apply Below
* Provides leadership to one or more teams designing and architecting infrastructure and service and provides input on best practices for reliability and functionality. Establishes direction to ensure accurate forecasting and ensure systems have adequate resources. Builds collaborative relationships with the software development team to create reliable, scalable infrastructures. Ensures alignment regarding data collection and contributes to standards for optimizing operations and infrastructure reliability.
Defines approaches for incident response activities to ensure service reliability. Ensures in-depth reports. Plays a key role in developing standards for identifying and recommending automation. Anticipates and explains the impact of changes, mentoring other managers on what to communicate. Defines approaches for escalating incidents and refines methods for documentation. Encourages experimenting with new technology, executing improvements, building site reliability knowledge, and providing clear data.
LI-ES2
** Responsibilities*
* ** Key Responsibilities*
* ** Capacity Ingestion and Management:*
* -Provides leadership for one or more teams designing and architecting infrastructure and/or service, providing input on the development of best practices for adhering to terms for reliability and functionality.
-Establishes direction for other managers and senior-level individuals to drive the forecasting of demands for infrastructure and respond to capacity needs, ensuring that systems have sufficient resources to meet current and future workloads and identifying and addressing resource gaps.
-Builds collaborative relationships with senior software development team members to design and develop infrastructures that are highly reliable and scalable, meeting stringent deployment requirements.
-Ensures teams align on expectations for identifying opportunities for prototyping and oversees prototyping initiatives (e.g., testing new applications or infrastructures, assisting in onboarding), experimenting with cutting-edge approaches.
** Incident and Service Lifecycle Management:*
* -Ensures alignment across teams regarding performing data collection, triage, technical analysis, and redirection, contributing to the development of standards to maintain and optimize operations and infrastructure reliability.
-Shares techniques across teams for monitoring of services, maintaining up-to-date knowledge of their performance, and thoroughly documenting their condition.
-Defines approaches for performing incident response, root cause analysis, and/or maintenance on assigned services (e.g., software installs, version upgrades, security updates, backup and recovery) and drives execution.
-Ensures teams provide in-depth health and performance reporting and coordinates managerial actions based on trends in data.
-Refines procedures for performing provisioning to support infrastructure, applications, and services, mentoring team members.
-Provides input on standards for decommissioning (e.g., shutting down servers, removing data from databases) to remove objects that are no longer needed.
** Automation:*
* -Plays a key role in developing standards for identifying and recommending opportunities for automation and reviewing potential benefits in terms of metrics across teams to ensure expectations are met.
-Ensures alignment on expectations for developing and drives the implementation of design, automation tools, or scripts.
-Refines strategies for conducting testing on highly complex automations to ensure they perform tasks correctly and produce expected results.
-Provides guidance and expertise to others testing automations.
** Technical Communication and Guidance:*
* -Shares expectations for release notes and communication of in-depth information about the scale, capacity, security, performance attributes, and requirements of services and technology with customers, cross-functional teams and leadership.
-Anticipates and explains the potential impact of infrastructure, feature, and tool changes, considering the strategic implications and goals.
-Takes a leadership role in mentoring other managers on what…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×