×
Register Here to Apply for Jobs or Post Jobs. X

Senior Elasticsearch Engineer

Remote / Online - Candidates ideally in
Phoenix, Maricopa County, Arizona, 85003, USA
Listing for: Chess.com
Remote/Work from Home position
Listed on 2026-07-24
Job specializations:
  • IT/Tech
    SRE/Site Reliability, Systems Engineer
Salary/Wage Range or Industry Benchmark: 140000 - 210000 USD Yearly USD 140000.00 210000.00 YEAR
Job Description & How to Apply Below

About You

is the world's largest chess platform with 235M+ members and ~20 million daily games. Our Elasticsearch and Open Search infrastructure underpins search, user activity, analytics, logging, and operational intelligence at massive scale, hundreds of terabytes across a dozen production clusters running on bare‑metal Kubernetes.

We're looking for a Senior Elasticsearch Engineer who can own the full lifecycle of our search and analytics data platform: capacity planning, cluster architecture, performance tuning, incident response, migration strategy, and operational excellence. You'll be the single point of deep expertise across all Elasticsearch and Open Search clusters at

This is not a monitoring‑from‑dashboards role. You'll be hands‑on with cluster internals, write ILM/ISM policies, push infrastructure changes through Git Ops, and make real‑time decisions about replica allocation when a cluster goes red.

What you’ll do Incident Response & Reliability
  • Shard allocation strategy for write‑heavy data streams at high throughput (millions of documents per minute)
  • Disk watermark management, retention policy tuning, and rollover orchestration for high‑volume indices
  • Performance optimization and I/O tuning on bare‑metal nodes
  • Write queue analysis, thread pool diagnostics, and shard rebalancing under load
  • Capacity planning and growth forecasting across clusters
  • On‑call ownership for Elasticsearch‑related incidents: cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades
  • Real‑time cluster triage and cross‑team coordination during production incidents
  • Post‑mortem authoring and systemic reliability improvements
  • Snapshot and disaster recovery management across clusters
Migration & Strategy
  • Elasticsearch‑to‑Open Search migration analysis and execution, including compatibility evaluation across ILM/ISM, security models, and plugin ecosystems
  • Version upgrade planning and rolling restart orchestration with zero‑downtime requirements
  • End‑to‑end new cluster provisioning and onboarding
Cross‑Team Enablement
  • Advise engineering teams on index design, mapping strategy, retention policies, and query optimization
  • Manage Kibana and Open Search Dashboards access and configuration for internal consumers
  • Define and maintain workload priority tiers across clusters
Preferred Skills
  • 7+ years operating Elasticsearch at scale (multi‑TB clusters, dozens of nodes, high write throughput)
  • Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management
  • Production experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator‑based deployments
  • Proficiency with Kubernetes operations for stateful workloads (Stateful Sets, persistent storage, resource management)
  • Hands‑on Linux systems administration with a focus on storage and I/O performance
  • Experience managing both Elasticsearch and Open Search in production, including an informed opinion on their respective trade‑offs
  • Incident command experience: ability to diagnose and mitigate cluster emergencies under pressure while communicating clearly to stakeholders
  • Git‑based infrastructure management (Git Ops):
    Helm charts, ArgoCD/Flux, infrastructure‑as‑code for cluster configuration
  • Fluency with the Elastic stack APIs: cluster administration, index templates, data streams, ILM policies, snapshot/restore
Bonus Experience
  • Open Search ISM policies and security plugin (fine‑grained access control)
  • GCS or S3 snapshot repository configuration and cross‑cluster replication
  • Grafana + Prometheus monitoring for Elasticsearch metrics
  • Kibana Discover, Dev Tools, and data view management at scale
  • Java internals relevant to Elasticsearch JVM tuning (heap sizing, GC tuning, circuit breakers)
  • Vault integration for secrets management in Kubernetes‑deployed search clusters
  • Fluentd/Fluent Bit log pipeline configuration feeding Open Search
  • Hardware selection experience for search‑optimized server configurations
  • Python or scripting for operational analysis and automation
What Makes This Role Special
  • Full autonomy. You are the Elasticsearch authority. You make the architecture calls, set the priorities, and own the outcomes.
  • Real scale. Hundreds of terabytes of data, billions of documents, millions of daily queries. The problems here don't exist at smaller companies.
  • Bare metal. No managed Elastic Cloud. You're operating directly on the hardware. This is hands‑on engineering.
  • Strategic impact. Your decisions on ES vs. Open Search migration, cluster topology, and capacity planning directly affect product capabilities and infrastructure costs.
  • Small team, big trust.  runs lean. You won't be buried in process or approvals. Ship changes, fix problems, improve systems.
About the Opportunity

This is a full‑time opportunity. We are 100% remote (work from anywhere!).

#J-18808-Ljbffr
Position Requirements
10+ Years work experience
To View & Apply for jobs on this site that accept applications from your location or country, tap the button below to make a Search.
(If this job is in fact in your jurisdiction, then you may be using a Proxy or VPN to access this site, and to progress further, you should change your connectivity to another mobile device or PC).
 
 
 
Search for further Jobs Here:
(Try combinations for better Results! Or enter less keywords for broader Results)
Location
Increase/decrease your Search Radius (miles)
0
200
Filters
Education Level
Experience Level (years)
Posted in last:
Salary