×
Register Here to Apply for Jobs or Post Jobs. X

Principal Site Reliability Engineer

Job in Nashville, Davidson County, Tennessee, 37230, USA
Listing for: Oracle
Full Time position
Listed on 2026-07-31
Job specializations:
  • IT/Tech
    Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability
Job Description & How to Apply Below
** Job Description*
* The successful candidate will serve as a senior technical authority, establish reliability standards, guide complex technical decisions, and lead improvements that reduce operational risk and manual effort. This individual must be comfortable moving between architecture and hands-on execution, including accessing deployed hosts, troubleshooting failed services, reviewing logs, correcting configurations, and validating production changes.

Designs and architects infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Exercises judgment when performing data collection to maintain and optimize operations and reliability. Leverages advanced knowledge to perform incident response and/or maintenance tasks. Provides comprehensive health and performance reporting. Identifies and recommends opportunities for automation.

Communicates comprehensive information about services and proactively anticipates and articulates the potential impact of changes. Provides comprehensive support for technology and documents incidents. Conducts advanced experiments with new tools and develops and maintains advanced knowledge of site reliability trends.

** Responsibilities*
* ** Key Responsibilities*
* + Design and architect reliable, secure, scalable, and maintainable infrastructure and services. Take proactive steps to ensure solutions meet defined reliability and functionality requirements.

+ Establish technical direction, engineering standards, and operational best practices across complex infrastructure and application environments.

+ Identify system dependencies, operational risks, capacity constraints, performance issues, and potential failure points before they affect service.

+ Translate business, client, security, and application requirements into practical infrastructure and reliability solutions.

+ Lead the installation, configuration, deployment, and validation of applications across Windows Server and Linux environments.

+ Oversee structured builds and deployments using runbooks, scripts, readiness assessments, change controls, and post-deployment validation.

+ Troubleshoot complex operating system, service, application, installation, patching, permissions, certificate, and connectivity issues.

+ Define and improve monitoring, alerting, logging, observability, capacity planning, and service-health practices.

+ Develop and promote automation that reduces manual effort, improves consistency, and lowers operational risk.

+ Own the administration, governance, and continuous improvement of an enterprise infrastructure automation platform, including maintaining automation standards, enabling engineering teams, and driving adoption of scalable, repeatable infrastructure management practices.

+ Lead operating system, middleware, and application patching initiatives, including change planning, rollback preparation, execution, and validation.

+ Direct major incident response, root cause analysis, corrective-action planning, and prevention of recurring failures.

+ Partner with cybersecurity teams on vulnerability remediation, system hardening, STIG compliance, and other security-driven changes.

+ Evaluate emerging technologies and recommend solutions that improve reliability, resilience, security, and operational efficiency.

+ Create and maintain technical standards, architecture documentation, runbooks, deployment procedures, and troubleshooting guides.

+ Provide technical leadership, mentorship, and design guidance to engineers across multiple teams.

+ Communicate technical risks, dependencies, decisions, and recommendations clearly to leadership and stakeholders.

** Core

Skills and Qualifications *
* ** Technical Leadership and Architecture*
* + Extensive experience in site reliability engineering, systems engineering, infrastructure architecture, production operations, or application hosting.

+ Demonstrated ability to design and support highly available, resilient, and secure enterprise systems.

+ Experience leading complex technical initiatives across…
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary