Site Reliability Engineer - Nashville, TN- Hybrid
Listed on 2026-07-01
-
IT/Tech
Systems Engineer, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Cybersecurity
Site Reliability Engineer Long Term Nashville, TN - Hybrid (3 Days In Office)
Provide SME level support for web technologies. Handle complex issues and problem management. Perform risk and security self-assessments across web infrastructure platforms. Leverage automation and development technologies to execute operational stability and maintenance tasks. Determine the reliability of our digital products, technology services, and the infrastructure that underpins them. Minimize the risk and impact of failures by engineering operational improvements, such as predictive monitoring, auto scaling or self-healing.
Respond to production incidents to gain first-hand experience of operational hotspots and to identify the root causes of problems. Collect and analyze operational data, define and monitor key metrics to identify and communicate areas for improvement. Apply a broad range of engineering practices with a focus on reliability, from instrumentation, performance analysis, and log analytics to automated testing, deployment, and operations.
Ensure the quality, security, reliability, and compliance of our solutions by applying our digital principles and implementing both functional and non-functional requirements.
Ideally 8+ years of experience in a similar position focused on web technologies. Bachelor/Master's Degree or equivalent focusing on information technology. Proficient in managing web hosting products such as Apache, Tomcat, IBM Web Sphere Application Server and IIS in a large and complex environment. Good knowledge in scripting and programming languages (Python/Java script/Shell/Power Shell). Knowledge in Chef, Puppet, Ansible, Kubernetes, IP Center or any automation technology.
Good understanding of infrastructure principles and system inter-dependencies outside the web hosting environment (e.g. networking, databases). Interested in learning new technologies and practices, reuse strategic platforms and standards, evaluate options, and make decisions with long-term sustainability in mind. Strong communicator, from making presentations to technical writing, fluent in English.
Knowledge on Azure Cloud or Pivotal Cloud Foundry platform. ITIL process know-how for incident, change and problem management.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).