More jobs:
Site Reliability Analyst
Job in
Saint Louis, St. Louis city, Missouri, 63101, USA
Listed on 2026-07-25
Listing for:
Apex Systems
Full Time
position Listed on 2026-07-25
Job specializations:
-
IT/Tech
IT Support, Cloud Computing: Infrastructure & Operations, SRE/Site Reliability, Systems Engineer
Job Description & How to Apply Below
Site Reliability Analyst
We are seeking a highly skilled Site Reliability Analyst to join our IT Service Delivery organization. This role is responsible for the end-to-end reliability, performance, observability, and operational excellence of a large-scale enterprise performance analytics platform. Initially, this position will act as a backup for system maintenance and support, transitioning long-term into a performance engineering role focused on troubleshooting complex system and application issues.
Key Responsibilities
- Support and maintain the Capacity Management (CPAC) environment, performing daily operational monitoring and system health checks.
- Troubleshoot performance issues, collector failures, missing data, and application stability concerns.
- Investigate system and application performance issues, assisting customers with root cause analysis.
- Design, build, and support a production monitoring and analytics platform in a Docker/Kubernetes containerized environment.
- Ensure high availability, reliability, and scalability of platform services while adhering to enterprise security and compliance standards.
- Develop operational runbooks, standard operating procedures, and support documentation.
- Perform manual OS patching, vulnerability remediation, and operate vulnerability scanning tools.
- Correlate logs, metrics, traces, and events using tools like Splunk, Prometheus, and Grafana to identify root causes.
- Analyze system resource utilization, including CPU, memory, disk I/O, and network latency.
- Lead technical discussions with support teams to investigate and resolve performance issues.
Required Qualifications
- 5+ years of Linux/Unix administration experience, including system build and deployment.
- 5+ years of experience supporting enterprise production applications.
- 3+ years of container platform administration experience (Docker/Kubernetes).
- Proven experience in troubleshooting performance issues in distributed production environments and performing root cause analysis.
- Experience with service maintenance, operational support, and performance management.
Technical Skills
- Programming and automation experience with one or more:
Python, No
SQL, Svelte, SQL, GIT, Bash/Shell scripting. - Experience with observability tools such as Splunk, Grafana, Prometheus, App Dynamics, or Open Telemetry.
- Familiarity with security scanning tools, vulnerability management, and patch management.
Professional Skills
- Advanced analytical problem-solving and critical thinking skills.
- A strong drive and attitude to learn, resolve issues, and not give up easily.
- Ability to communicate recommendations effectively to both technical and non-technical audiences.
- Must be a U.S. Person.
Preferred Qualifications
- Exposure to Airflow.
- Site Reliability Engineering (SRE) experience.
- Background in performance engineering or capacity planning/management.
- Cloud platform experience (AWS, Azure, GCP).
- Certification in Kubernetes or Docker.
- Broad technical exposure across various OS platforms (e.g., HPUX, AIX, VMware), networking protocols, and storage technologies.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×