Senior Site Reliability Engineer
Listed on 2026-08-15
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Administrator, IT Infrastructure, Systems Engineer
Job Title
Infrastructure and Service Architect
Job DescriptionTakes proactive steps to design and architect infrastructure and service to ensure reliability and functionality. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Performs data collection to maintain and optimize operations and reliability. Leverages knowledge to perform incident response and/or maintenance tasks. Provides health and performance reporting. Identifies opportunities for automation.
Communicates about services and identifies and explains the potential impact of changes. Provides support for technology and documents incidents. Experiments with new tools and assesses potential impact and develops knowledge of site reliability trends.
Design and architect reliable, secure, and maintainable infrastructure and services. Take proactive steps to ensure solutions meet defined reliability and functionality requirements.
- Identify operational risks, dependencies, performance issues, and potential failure points before they affect service.
- Translate business and application requirements into practical technical solutions.
- Install, configure, deploy, and validate applications across Windows Server and Linux environments.
- Perform structured builds and deployments using runbooks, scripts, readiness checks, and post-build validation.
- Troubleshoot operating system, service, application, installation, patching, permissions, certificate, and connectivity issues.
- Monitor system performance and implement improvements to availability, reliability, and operational efficiency.
- Develop and maintain scripts and automation that reduce manual effort and improve consistency.
- Plan and execute operating system, middleware, and application patching using established change-control and rollback procedures.
- Lead or support incident investigations, root cause analyses, and corrective actions.
- Support vulnerability remediation, system hardening, STIG compliance, and other security-driven changes.
- Maintain accurate runbooks, deployment procedures, troubleshooting guides, and operational records.
- Provide technical guidance and mentorship to junior engineers.
- Communicate status, risks, blockers, and escalation details clearly to stakeholders.
Skills and Qualifications
Windows and Linux System Administration
- Hands-on experience administering Windows Server and/or Linux systems.
- Ability to access deployed hosts and perform post-deployment configuration and validation.
- Experience installing, configuring, and validating applications in Windows Server and Linux environments.
- Ability to troubleshoot operating-system-level, service-level, and application-level issues.
- Working knowledge of system services, permissions, configuration files, logs, and resource utilization.
Manual Build and Deployment Experience
- Experience performing structured build and deployment tasks using runbooks, deployment guides, and technical procedures.
- Ability to execute scripts, validate outputs, and correct common build or configuration issues.
- Familiarity with build handoffs, environment-readiness checks, deployment validation, and post-build verification.
- Ability to follow detailed implementation steps while identifying and documenting exceptions or deviations.
Troubleshooting and Operational Support
- Ability to investigate failed services, installation errors, patching failures, application startup problems, permissions issues, and connectivity failures.
- Experience reviewing logs, event viewers, service status, configuration files, ports, certificates.
- Strong analytical and problem-solving skills, with the ability to isolate root causes and recommend practical solutions.
- Ability to escalate issues clearly by documenting symptoms, troubleshooting steps, findings, impact, and recommended actions.
- Experience supporting production or other business-critical environments.
Scripting and Automation
Hands-on experience with one or more of the following:
- Power Shell
- Bash
- Python
- Ansible
- Chef
Candidates should be able to run, modify, validate, and troubleshoot existing scripts and understand basic automation concepts.
Patching and Software Maintenance
- Experience applying operating system, middleware, and application patches.
- Ability to follow patching procedures, validate successful completion, and troubleshoot failures.
- Understanding of maintenance windows, change control, rollback planning, and post-change validation.
Cloud and OCI Familiarity
- Familiarity with Oracle Cloud Infrastructure or another major cloud platform.
- Understanding of cloud compute, storage, networking, identity, and environment-provisioning concepts.
- Experience supporting applications in cloud-hosted or hybrid environments.
Network Troubleshooting
- Working knowledge of DNS, firewalls, routing, load balancers, ports, and certificates.
- Ability to identify basic connectivity issues between hosts, applications, and services.
- Familiarity with standard network diagnostic…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).