Senior Grafana Engineer
Job in
Atlanta, Fulton County, Georgia, 30383, USA
Listed on 2026-07-23
Listing for:
Capgemini America, Inc.
Full Time
position Listed on 2026-07-23
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, AWS, SRE/Site Reliability
Job Description & How to Apply Below
Location
Hybrid - Preferred Troy, NY or Chicago, Atlanta, Nashville, Dallas, NJ
About the job you're considering
At Capgemini, you will collaborate with cross-functional teams to deliver innovative technology solutions that drive business value and enhance client experiences. You will contribute to the design, development, and continuous improvement of scalable, high-quality solutions in a dynamic and collaborative environment.
Executive Summary
The Senior Grafana Engineer will be responsible for enterprise observability platform engineering, Grafana Cloud migration, Open Telemetry adoption, automation, integrations, production support, and operational excellence across AWS and hybrid environments.
Core Responsibilities
* Lead Grafana Cloud platform engineering and migration from New Relic.
* Design enterprise observability architecture using Grafana Alloy, Open Telemetry, Mimir, Loki, Tempo, and IRM.
* Develop dashboards, alerts, service health views, SLIs, SLOs, and operational reporting.
* Support production readiness, incident management, MTTR reduction, and operational excellence.
* Create runbooks, onboarding standards, and reusable observability templates.
Grafana Expertise
Required Skills
* Grafana Cloud
* Grafana Scenes
* PromQL
* LogQL
* TraceQL
* Executive and technical dashboard design (UX)
* Grafana Alloy
* Open Telemetry instrumentation patterns
Responsibilities
* Build and maintain enterprise-level observability solutions.
* Design dashboards for both executive leadership and technical teams.
* Implement monitoring, logging, tracing, and incident response capabilities.
AWS & Platform Engineering
Required Experience
AWS
* ECS
* Fargate
* EC2
* Lambda
* Cloud Watch
Additional Skills
* Hybrid Cloud and On-Premises Observability
* AWS Systems Manager (SSM)
* Infrastructure as Code (Terraform)
* CI/CD Automation
Responsibilities
* Monitor and support cloud-native and hybrid workloads.
* Automate observability deployment and management using Terraform and CI/CD pipelines.
* Integrate AWS telemetry into Grafana Cloud.
Integration Experience
ITSM & Operations Integrations
* Service Now ITSM Integration
* Service Now CMDB Integration and Service Mapping
* Pager Duty Incident Management Integration
* Slack and Chat Ops Integrations
Analytics & Data Integrations
* AWS Cost Explorer and Fin Ops Dashboards
* Snowflake Analytics Integration
* dbt Data Model Integration
* REST APIs and Webhooks
* Enterprise Platform Integrations
Automation & AI
Required Skills
* Python Development and Automation
* Data Pipelines and Telemetry Enrichment
* Model Context Protocol (MCP)
* LLM Integrations
* Agentic Workflows
* AI-Assisted Incident Analysis
* Operational Intelligence Solutions
Responsibilities
* Automate operational processes and observability workflows.
* Build AI-powered solutions for incident analysis and troubleshooting.
* Enhance telemetry data with enrichment pipelines and intelligent insights.
Security
Required Experience
* Cyber Ark Integration
* Secrets Management
* IAM Governance
* Secure Observability Platform Administration
Responsibilities
* Ensure secure access management and credential handling.
* Implement governance and compliance controls across observability platforms.
Ideal Candidate Profile
Technical Expertise
* Grafana Cloud Platform Engineering
* Open Telemetry Architecture
* AWS Infrastructure Monitoring
* Terraform & Automation
* Python Development
* Service Now Integrations
* Incident Management & SRE Practices
Desired Outcomes
* Successfully migrate from New Relic to Grafana Cloud.
* Establish enterprise observability standards.
* Improve service reliability and reduce MTTR.
* Deliver scalable monitoring, logging, tracing, and alerting solutions.
* Enable AI-driven operational intelligence capabilities.
Key Technologies
Observability
* Grafana Cloud
* Grafana Alloy
* Mimir
* Loki
* Tempo
* IRM
* Open Telemetry
Query Languages
* PromQL
* LogQL
* TraceQL
Cloud & Infrastructure
* AWS
* ECS
* Fargate
* EC2
* Lambda
* Cloud Watch
* SSM
* Terraform
Integrations
* Service Now
* Pager Duty
* Slack
* Snowflake
* dbt
* REST APIs
* Webhooks
Automation & AI
* Python
* MCP
* LLMs
* Agentic Workflows
Security
* Cyber Ark
* IAM
* Secrets Management
Primary Focus:
Enterprise Observability Engineering, Grafana Cloud Migration, Open Telemetry Adoption, AWS Platform Monitoring, Automation, Incident Management, and AI-Enabled Operations.
#LI-SS1
Position Requirements
10+ Years
work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×