Principal Site Reliability Engineer
Listed on 2026-08-01
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, Cybersecurity, SRE/Site Reliability
Job Description
This role combines strategic architecture with practical systems engineering, deployment, automation, patching, troubleshooting, incident response, and compliance support. The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud Infrastructure, hybrid cloud, and legacy environments while partnering with engineering, operations, cybersecurity, networking, application, and client-facing teams.
Job DescriptionThis role combines strategic architecture with practical systems engineering, deployment, automation, patching, troubleshooting, incident response, and compliance support. The Principal Site Reliability Engineer will work across Windows, Linux, Oracle Cloud Infrastructure, hybrid cloud, and legacy environments while partnering with engineering, operations, cybersecurity, networking, application, and client-facing teams.
The successful candidate will serve as a senior technical authority, establish reliability standards, guide complex technical decisions, and lead improvements that reduce operational risk and manual effort. This individual must be comfortable moving between architecture and hands-on execution, including accessing deployed hosts, troubleshooting failed services, reviewing logs, correcting configurations, and validating production changes.
ResponsibilitiesKey Responsibilities
- Design and architect reliable, secure, scalable, and maintainable infrastructure and services. Take proactive steps to ensure solutions meet defined reliability and functionality requirements.
- Establish technical direction, engineering standards, and operational best practices across complex infrastructure and application environments.
- Identify system dependencies, operational risks, capacity constraints, performance issues, and potential failure points before they affect service.
- Translate business, client, security, and application requirements into practical infrastructure and reliability solutions.
- Lead the installation, configuration, deployment, and validation of applications across Windows Server and Linux environments.
- Oversee structured builds and deployments using runbooks, scripts, readiness assessments, change controls, and post-deployment validation.
- Troubleshoot complex operating system, service, application, installation, patching, permissions, certificate, and connectivity issues.
- Define and improve monitoring, alerting, logging, observability, capacity planning, and service-health practices.
- Develop and promote automation that reduces manual effort, improves consistency, and lowers operational risk.
- Lead operating system, middleware, and application patching initiatives, including change planning, rollback preparation, execution, and validation.
- Direct major incident response, root cause analysis, corrective-action planning, and prevention of recurring failures.
- Partner with cybersecurity teams on vulnerability remediation, system hardening, STIG compliance, and other security-driven changes.
- Evaluate emerging technologies and recommend solutions that improve reliability, resilience, security, and operational efficiency.
- Create and maintain technical standards, architecture documentation, runbooks, deployment procedures, and troubleshooting guides.
- Provide technical leadership, mentorship, and design guidance to engineers across multiple teams.
- Communicate technical risks, dependencies, decisions, and recommendations clearly to leadership and stakeholders.
- Extensive experience in site reliability engineering, systems engineering, infrastructure architecture, production operations, or application hosting.
- Demonstrated ability to design and support highly available, resilient, and secure enterprise systems.
- Experience leading complex technical initiatives across engineering, operations, security, networking, and application teams.
- Ability to make sound architectural decisions, evaluate tradeoffs, and communicate recommendations to technical and nontechnical stakeholders.
- Experience defining engineering standards, operational controls, and reliability practices.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).