Senior Site Reliability Engineer
Listed on 2026-07-31
-
IT/Tech
Systems Administrator, Systems Engineer, Cloud Computing: Infrastructure & Operations, IT Infrastructure
Infrastructure And Service Architect
Takes proactive steps to design and architect infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Performs data collection to maintain and optimize operations and reliability. Leverages knowledge to perform incident response and/or maintenance tasks. Provides health and performance reporting. Identifies opportunities for automation.
Communicates about services and identifies and explains the potential impact of changes. Provides support for technology and documents incidents. Experiments with new tools and assesses potential impact and develops knowledge of site reliability trends.
Key Responsibilities
- Design and architect reliable, secure, and maintainable infrastructure and services. Take proactive steps to ensure solutions meet defined reliability and functionality requirements.
- Identify operational risks, dependencies, performance issues, and potential failure points before they affect service.
- Translate business and application requirements into practical technical solutions.
- Install, configure, deploy, and validate applications across Windows Server and Linux environments.
- Perform structured builds and deployments using runbooks, scripts, readiness checks, and post-build validation.
- Troubleshoot operating system, service, application, installation, patching, permissions, certificate, and connectivity issues.
- Monitor system performance and implement improvements to availability, reliability, and operational efficiency.
- Develop and maintain scripts and automation that reduce manual effort and improve consistency.
- Plan and execute operating system, middleware, and application patching using established change-control and rollback procedures.
- Lead or support incident investigations, root cause analyses, and corrective actions.
- Support vulnerability remediation, system hardening, STIG compliance, and other security-driven changes.
- Maintain accurate runbooks, deployment procedures, troubleshooting guides, and operational records.
- Provide technical guidance and mentorship to junior engineers.
- Communicate status, risks, blockers, and escalation details clearly to stakeholders.
Core Skills and Qualifications
Windows and Linux System Administration
- Hands-on experience administering Windows Server and/or Linux systems.
- Ability to access deployed hosts and perform post-deployment configuration and validation.
- Experience installing, configuring, and validating applications in Windows Server and Linux environments.
- Ability to troubleshoot operating-system-level, service-level, and application-level issues.
- Working knowledge of system services, permissions, configuration files, logs, and resource utilization.
Manual Build and Deployment Experience
- Experience performing structured build and deployment tasks using runbooks, deployment guides, and technical procedures.
- Ability to execute scripts, validate outputs, and correct common build or configuration issues.
- Familiarity with build handoffs, environment-readiness checks, deployment validation, and post-build verification.
- Ability to follow detailed implementation steps while identifying and documenting exceptions or deviations.
Troubleshooting and Operational Support
- Ability to investigate failed services, installation errors, patching failures, application startup problems, permissions issues, and connectivity failures.
- Experience reviewing logs, event viewers, service status, configuration files, ports, certificates, and permissions.
- Strong analytical and problem-solving skills, with the ability to isolate root causes and recommend practical solutions.
- Ability to escalate issues clearly by documenting symptoms, troubleshooting steps, findings, impact, and recommended actions.
- Experience supporting production or other business-critical environments.
Scripting and Automation
Hands-on experience with one or more of the following:
- Power Shell
- Bash
- Python
- Ansible
- Chef
Candidates should be able to run, modify, validate, and troubleshoot existing scripts and understand basic automation concepts.
Patc…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).