Remote Site Reliability Engineer
Listed on 2026-08-05
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Data Engineering
Site Reliability Engineering Team
Overview:
Surrounding American Family mergers and acquisitions. Currently have 15 separate data centers with a variety of technical stacks that need to be consolidated into 3 data centers (in Ashburn, Dallas, and Seattle). Squad as a whole is responsible for the core platform services, split into 2 teams: 1 is Platform Engineering, and the other is Site Reliability Engineering. This is for the Site Reliability Engineering team.
Need in-depth knowledge of Linux, Windows Administration, setting up racks, servers
Infrastructure as code automations in Terraform, Puppet to automate aspects migrations
Hybrid cloud environment, do have AWS and building a private cloud
Different data centers are currently running on different infrastructure, need to bring under 1 central stack: consisting of HPE Performance Cluster, Windows (older and newer versions), Linux (older and newer versions), JBOSS, Tomcat servers
Terraform IaC, Octopus, Jenkins, Puppet
Day to day: incident management, looking at alerts (Dynatrace, Datadog), certificate management, setting up the core systems--CPU, memory, storage, adjustments/configurations. Where can we automate?
2 profiles: 1 more focused on building the data centers; 1 more on the SRE side, supporting current systems that are already out there. More of a Systems Engineer with Data Center focus vs SRE.
Required Skills:
Linux
Windows administration
Terraform
Puppet or similar
Scripting:
Python, Shell, BASH
Public or Private Cloud experience (Private preferred)
Nice to Have:
Data center consolidation experience
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).