Grafana Engineer; Observability & Monitoring
Listed on 2026-09-03
-
IT/Tech
SRE/Site Reliability, Cloud Computing: Infrastructure & Operations, Systems Engineer, IT Infrastructure
Purple Drive
Grafana Engineer – Observability & Monitoring
We are seeking a skilled Grafana Engineer with 4–6 years of experience in designing, implementing, and maintaining enterprise monitoring and observability solutions. The ideal candidate will have hands-on expertise in Grafana, Prometheus, Loki, Open Telemetry, Splunk, Kubernetes, and cloud monitoring platforms.
The role involves building dashboards, configuring alerts, integrating multiple monitoring data sources, implementing observability strategies, automating monitoring infrastructure, and collaborating with Dev Ops, SRE, Infrastructure, and Application Support teams to ensure the reliability, availability, and performance of enterprise applications.
Experience
Required:
- 4–6 Years Monitoring & Observability Engineering
- Site Reliability Engineering (SRE)
- Dev Ops / Platform Engineering
- Design, develop, and maintain Grafana dashboards and visualizations.
- Build operational, infrastructure, application, and business monitoring dashboards.
- Configure Grafana Alerting, notification channels, and alert rules.
- Administer and maintain Grafana Enterprise environments.
- Optimize dashboards for usability and performance.
- Design enterprise observability solutions covering:
Metrics, Logs, Distributed Traces - Implement monitoring strategies for:
Kubernetes clusters, Containers, Virtual Machines, Cloud Infrastructure, Enterprise Applications - Configure proactive monitoring and alerting to reduce downtime.
- Perform root cause analysis using monitoring and log data.
Integrate Grafana with enterprise monitoring platforms including:
- Prometheus
- Loki
- Tempo
- Open Telemetry
- Elasticsearch / Open Search
- InfluxDB
- SQL Databases
- Splunk
- AWS Cloud Watch
- Azure Monitor
- Google Cloud Operations (GCP)
- Implement cloud-native monitoring solutions on: AWS, Microsoft Azure, Google Cloud Platform (GCP)
- Monitor containerized workloads and Kubernetes clusters.
- Support enterprise platform observability.
- Automate dashboard deployment using Infrastructure-as-Code (IaC).
- Implement CI/CD pipelines for monitoring configurations.
- Use:
Terraform, Ansible - Develop automation scripts using:
Python, Bash, Power Shell
- Tune monitoring systems to reduce alert fatigue.
- Perform capacity planning.
- Conduct system performance analysis.
- Optimize monitoring configurations for scalability.
- Create monitoring standards and operational runbooks.
- Maintain technical documentation.
- Ensure monitoring platform security and governance.
- Implement access control and monitoring best practices.
- Grafana
- Grafana Enterprise
- Grafana Alerting
- Prometheus
- Loki
- Tempo
- Open Telemetry
- PromQL
- LogQL
- Splunk
- Elasticsearch
- Open Search
- InfluxDB
- AWS Cloud Watch
- Azure Monitor
- Google Cloud Operations (GCP)
- Kubernetes
- Docker
- CI/CD
- Terraform
- Ansible
- Python
- Bash
- Shell Scripting
- Power Shell
- Metrics
- Logs
- Distributed Tracing
- Root Cause Analysis (RCA)
- Incident Management
- Performance Monitoring
- Grafana Enterprise
- Open Telemetry
- AIOps
- Dynatrace
- Datadog
- App Dynamics
- New Relic
- Prometheus Federation
- Service Mesh Monitoring
- Strong analytical and troubleshooting skills
- Excellent communication skills
- Stakeholder management
- Problem-solving mindset
- Team collaboration
- Documentation and knowledge sharing
- Ability to work independently
- Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).