Senior Site Reliability Engineer
Listed on 2026-08-13
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer
Senior Site Reliability Engineer
The expertise you have:
Understanding of critical production applications and their infrastructure dependencies to supervise, analyze and discover anomalies. Deep knowledge of application request rates, transactions per second, client volumes to plan for adequate capacity to handle future growth. Knowledge of SLI/SLO/SLA definitions and golden signals for critical applications. Understanding of telemetry, anomaly detection and observability. Skills in end-to-end development of reports and interactive dashboards with impactful visualizations in BI Publisher and/or Tableau.
Experience with monitoring tools (Datadog, Splunk, Dynatrace, Team Quest, Metric beat, ELK, Prometheus, Grafana etc). Good knowledge of virtualized environment operations characteristics – VMware/vSphere, Open Stack. Solid grasp of modern application configurations and frameworks. Exploring new trends in automation area and deploying in day-to-day activities. In-depth understanding of application modernization; microservices, API’s, etc. Shown knowledge of AWS and Azure cloud compute and storage services.
AWS certification is a plus. Keen focus on process to improve performance, presenting innovative recommendations, and implementing resolutions across the team. Able to explain complicated or technical information in a simple way to non-technical audiences. Experience as a developer (e.g.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).