More jobs:
Elastic Observability Architect
Job in
Kankakee, Kankakee County, Illinois, 60901, USA
Listed on 2026-08-05
Listing for:
Argyle Infotech
Full Time
position Listed on 2026-08-05
Job specializations:
-
IT/Tech
Cloud Computing: Infrastructure & Operations, Systems Engineer, SRE/Site Reliability, Data Engineering
Job Description & How to Apply Below
Elastic Observability Architect
We are seeking an experienced Elastic Observability Architect to lead the design, implementation, optimization, and governance of enterprise-scale observability solutions utilizing the Elastic Stack (ELK). The ideal candidate will possess deep expertise in Elasticsearch architecture, large-scale log analytics, performance optimization, distributed systems, and cloud-native observability platforms. This role requires a hands-on architect who can drive technical strategy, mentor engineering teams, and ensure the reliability, scalability, and security of enterprise observability environments supporting mission-critical applications and infrastructure.
Key Responsibilities- Architect, design, and manage enterprise-scale Elastic Stack (ELK) environments.
- Develop scalable and highly available observability solutions for logs, metrics, traces, and application monitoring.
- Design distributed Elasticsearch clusters with high availability, fault tolerance, and disaster recovery capabilities.
- Establish architecture standards, best practices, and governance frameworks for Elastic implementations.
- Configure and manage Elasticsearch clusters across production environments.
- Design and optimize indexing strategies, shard allocation, mappings, and query performance.
- Implement Index Lifecycle Management (ILM), Data Streams, and retention policies.
- Troubleshoot cluster health issues, garbage collection bottlenecks, indexing latency, and query performance challenges.
- Perform capacity planning and infrastructure optimization.
- Design and maintain robust ingestion pipelines using Logstash, Beats, Kafka, and related technologies.
- Integrate diverse enterprise data sources into the Elastic ecosystem.
- Ensure reliable, scalable, and secure data collection and processing workflows.
- Develop advanced Kibana dashboards, visualizations, alerts, and executive reporting solutions.
- Implement enterprise monitoring and observability strategies for infrastructure and applications.
- Deliver actionable operational insights through real-time analytics and reporting.
- Implement role-based access control (RBAC), TLS encryption, and Elastic security best practices.
- Ensure compliance with organizational security standards and governance policies.
- Manage authentication, authorization, and audit requirements.
- Provide technical leadership, architectural guidance, and design reviews.
- Mentor engineers and administrators on Elastic Stack best practices.
- Collaborate with infrastructure, cloud, Dev Ops, SRE, and application teams to drive observability initiatives.
- Lead troubleshooting efforts for critical production incidents.
- 10+ years of overall IT experience with strong expertise in observability and monitoring platforms.
- 5+ years of hands-on experience architecting and administering Elasticsearch/Elastic Stack environments.
- Expert-level knowledge of:
- Elasticsearch Cluster Architecture
- Distributed Systems Design
- Indexing, Sharding, and Query DSL
- Performance Tuning and Optimization
- Large-Scale Log and Metrics Analytics
- High Availability (HA) and Disaster Recovery (DR)
- Index Lifecycle Management (ILM)
- Data Streams
- Production Troubleshooting and Root Cause Analysis
- Elastic Stack
- Elasticsearch (Expert)
- Kibana
- Logstash
- Elastic Agent / Beats
- Elastic Security
- Elastic Observability
- Data Streaming & Integration
- Kafka
- Fluentd / Fluent Bit
- Data Ingestion Frameworks
- Cloud Platforms
- AWS
- Azure
- GCP
- Containerization & Orchestration
- Docker
- Kubernetes
- Dev Ops & Automation
- Terraform
- Ansible
- CI/CD Pipelines
- Infrastructure as Code (IaC)
- Security
- RBAC
- TLS/SSL
- Identity and Access Management
- Programming & Scripting
- Python
- Java
- Shell Scripting
- Experience implementing enterprise observability solutions for large-scale distributed environments.
- Background in Site Reliability Engineering (SRE) and platform engineering.
- Experience with Open Telemetry and modern observability frameworks.
- Elastic Certified Engineer or related certifications.
- Strong understanding of cloud-native architectures and microservices.
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
Search for further Jobs Here:
×