Senior Splunk & Observability Engineer
Listed on 2026-09-14
-
IT/Tech
AWS, SRE/Site Reliability, Cybersecurity
Job Title:
Senior Splunk & Observability Engineer
Job Type: Permanent Full Time
Position Description
Seeking an experienced Senior Splunk & Observability Engineer to support enterprise scale observability platforms within a complex, AWS based production environment. This is a senior hands on role combining advanced Splunk administration and engineering, AWS production support, observability, and automation.
The engineer will manage and support Splunk Enterprise running on AWS, working across the full data lifecycle from source onboarding and ingestion through indexing, search, dashboards, alerts, retention, and recovery. Responsibilities include developing and optimizing SPL queries, troubleshooting ingestion and performance issues, improving platform stability, supporting upgrades and vulnerability remediation, and restoring or replaying data following service interruptions.
The role will also work extensively with AWS, Amazon Cloud Watch, Open Telemetry, and Python based automation to improve enterprise monitoring and resiliency. The engineer will build and maintain Open Telemetry agents and collectors, support observability integrations across technology stacks, and help automate platform operations and vulnerability remediation.
This position is required in one of the following locations:
Lafayette, LA, Knoxville, TN, Columbia, SC, Birmingham, AL
- Administer, engineer, upgrade, and support Splunk Enterprise in AWS.
- Deploy, configure, and troubleshoot Splunk components hosted on Amazon EC2.
- Manage the end to end Splunk data lifecycle, including ingestion, indexing, search, dashboards, alerts, retention, and recovery.
- Develop and optimize complex SPL queries and troubleshoot query performance issues.
- Build and maintain enterprise dashboards for production monitoring, resiliency, incident response, and outage triage.
- Troubleshoot ingestion failures, problematic indexes, missing or delayed data, unexpected data volume growth, and platform availability issues.
- Restore data flow and replay data following ingestion or platform interruptions.
- Optimize data ingestion and retention to reduce noise, storage consumption, and operating costs.
- Integrate Splunk, Amazon Cloud Watch, and Open Telemetry for logs, metrics, dashboards, and alerts.
- Develop Python based automation for observability and platform operations.
- Build and maintain Open Telemetry agent and collector packages across multiple technology platforms.
- Support Splunk upgrades, patching, configuration changes, plugins, and vulnerability remediation.
- Use Git Lab, Git Hub, Terraform Enterprise, JSON, and CI/CD tooling to support infrastructure, configuration, automation, and package deployments.
- Lead technical troubleshooting calls and coordinate incident investigation and remediation across multiple teams.
- Participate in an on call rotation and provide after hours support when required for production incidents, upgrades, and planned changes.
- 5+ years of hands on Splunk Enterprise administration and engineering experience in large production environments.
- Strong understanding of the Splunk ecosystem, including data ingestion, forwarders, indexers, search heads, indexes, dashboards, alerts, and recovery.
- Advanced SPL query development and performance optimization skills.
- Hands on experience deploying, configuring, upgrading, troubleshooting, and supporting Splunk on AWS.
- Strong AWS production experience, particularly with Amazon EC2, Application Load Balancer, and Amazon Cloud Watch.
- Experience building complex Splunk dashboards for production monitoring, resiliency, incident triage, and operational use cases.
- Strong understanding of Open Telemetry, including agents,…
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).