Elasticsearch on Kubernetes
Elasticsearch on Kubernetes
In short: Running Elasticsearch on Kubernetes means using the official ECK (Elastic Cloud on Kubernetes) operator to deploy master, data, ingest, and coordinating node roles as separate StatefulSets, each backed by PersistentVolumeClaims on SSD StorageClasses. ECK automates TLS, autoscaling, snapshot lifecycle backups, and zero-downtime rolling upgrades while preserving cluster health.
Deploy and operate production-grade Elasticsearch clusters on Kubernetes with the ECK operator, auto-scaling, rolling upgrades, and enterprise-grade search infrastructure.
JusDB engineers production-grade Elasticsearch on Kubernetes deployments leveraging the official Elastic Cloud on Kubernetes (ECK) operator architecture. Certified DBREs isolate dedicated master node quorums, configure high-performance local NVMe storage classes, enforce zone-aware shard allocation, and orchestrate zero-downtime rolling stateful upgrades, backed by contractual 15-minute Sev-1 response SLAs and SOC 2 Type II compliance.
What we deliver
Comprehensive Elasticsearch on Kubernetes Services
From ECK operator deployment to production monitoring, we provide end-to-end Elasticsearch on Kubernetes solutions for search-intensive workloads.
ECK Operator Deployment
Deploy and configure the official Elastic Cloud on Kubernetes operator for automated cluster lifecycle management
- ECK operator installation and CRD setup
- Custom resource definitions for Elasticsearch
- TLS certificate auto-management
- Secure settings via Kubernetes Secrets
Helm Chart Management
Production-ready Helm chart configuration and management for repeatable, version-controlled deployments
- Custom values.yaml for each environment
- Helm release lifecycle management
- Chart versioning and rollback strategy
- GitOps integration (ArgoCD / Flux)
Cluster Topology on K8s
Design and deploy dedicated master, data, ingest, and coordinating node pools with proper resource isolation
- Dedicated master-eligible node StatefulSets
- Hot-warm-cold data tier architecture
- Ingest node pipeline configuration
- Coordinating-only nodes for search routing
Storage & Persistence
Configure persistent volumes, storage classes, and data durability for Elasticsearch on Kubernetes
- PVC-backed StatefulSet storage
- SSD StorageClass selection and tuning
- Volume expansion and resize policies
- Local PV vs network-attached storage
Monitoring & Observability
Full-stack monitoring with Kibana on Kubernetes, Prometheus exporters, and Grafana dashboards
- Kibana deployment and configuration on K8s
- Prometheus Elasticsearch Exporter
- Grafana dashboards for cluster health
- Alerting on shard allocation and node status
Backup & Snapshots
Automated snapshot lifecycle management with S3, GCS, or Azure Blob storage for disaster recovery
- Snapshot lifecycle management (SLM) policies
- S3 / GCS / Azure Blob snapshot repositories
- Automated backup scheduling and retention
- Cross-cluster snapshot restore and DR drills
Why Kubernetes
Why Run Elasticsearch on Kubernetes?
Kubernetes provides the orchestration layer that Elasticsearch needs for automated scaling, self-healing, and declarative cluster management in production environments.
Automated Scaling
Scale data, ingest, and coordinating nodes independently based on workload demands. ECK autoscaling policies combine with Kubernetes HPA/VPA to automatically add or remove Elasticsearch pods as indexing or search traffic fluctuates.
Zero-Downtime Rolling Upgrades
ECK orchestrates rolling upgrades one node at a time, handling shard migration, health checks, and version compatibility validation automatically. Your cluster stays available throughout the entire upgrade process.
Self-Healing Infrastructure
Kubernetes automatically restarts failed Elasticsearch pods, reschedules them to healthy nodes, and maintains the desired replica count. Combined with Elasticsearch's shard replication, this delivers robust fault tolerance.
Declarative Cluster Management
Define your entire Elasticsearch topology as Kubernetes custom resources. Version-control your cluster configuration, enable GitOps workflows, and reproduce identical environments across dev, staging, and production.
Elasticsearch on K8s Key Metrics
Node roles
ECK Architecture on Kubernetes
Understanding the node roles and Kubernetes resources that make up a production Elasticsearch deployment on K8s.
Master Nodes
StatefulSet (3 replicas)Dedicated master-eligible nodes handle cluster state, shard allocation, and index metadata management. Run as a 3-node StatefulSet for quorum.
- Cluster state management
- Shard allocation decisions
- Index creation / deletion
- Lightweight resource footprint
Data Nodes
StatefulSet (scalable)Data nodes store index shards and execute search and indexing operations. Sized with high storage and memory, scaled horizontally based on data volume.
- Hot / warm / cold tiering
- PVC-backed persistent storage
- CPU and memory intensive
- Horizontal auto-scaling
Ingest Nodes
StatefulSet (scalable)Ingest nodes run preprocessing pipelines (grok, dissect, enrichment) before documents are indexed. Isolate pipeline load from data node resources.
- Ingest pipeline execution
- Document enrichment
- Grok / dissect parsing
- Independent scaling
Coordinating Nodes
StatefulSet (scalable)Coordinating-only nodes act as smart load balancers, routing search requests and aggregating results across data nodes without holding data.
- Search request routing
- Scatter-gather aggregation
- Client-facing endpoints
- Reduce data node load
Methodology
Our Elasticsearch on Kubernetes Implementation Process
A proven methodology for deploying production-ready Elasticsearch clusters on Kubernetes with comprehensive testing and validation.
Architecture & Planning
Analyze search and indexing workloads, data volume, and query patterns. Design cluster topology, node roles, storage classes, and resource quotas for your Kubernetes environment.
ECK Deployment
Deploy the ECK operator, configure Elasticsearch custom resources, set up TLS, Kubernetes Secrets for credentials, and provision PersistentVolumes with appropriate StorageClasses.
Integration & Testing
Integrate with ingestion pipelines (Logstash, Beats, Filebeat), deploy Kibana, configure index templates, run failover tests, and benchmark indexing throughput and query latency.
Production & Operations
Go live with monitoring, alerting, snapshot policies, and auto-scaling. Provide runbooks, on-call playbooks, and ongoing support for rolling upgrades and capacity planning.
How JusDB DBRE Kubernetes Engineering compares to alternative paths.
Operating stateful Elasticsearch clusters on Kubernetes without dedicated DBRE operator management risks master quorum loss, PVC volume attachment deadlocks, and cascading shard rebalance storms during node drains. Here is how our certified cloud-native engineering compares across core evaluation vectors:
| Evaluation Vector | JusDB DBRE | In-House DBA | Legacy Agency | Developer Generalist |
|---|---|---|---|---|
| Elastic Cloud on Kubernetes (ECK) Operator Architecture | Deploys the official ECK operator with multi-NodeSet CustomResourceDefinitions, separating dedicated master, data (hot/warm/cold), and ingest roles with declarative GitOps reconcilers. | Uses unmaintained community Helm charts or monolithic manifests without operator reconciliation, requiring manual kubectl interventions during rolling mutations. | Deploys naked Kubernetes StatefulSets without operator intelligence, breaking elasticsearch.yml configuration discovery and internal TLS rotation. | Deploys Elasticsearch as a standalone single-replica Deployment without StatefulSet semantics, leading to pod restart data loss. |
| Dedicated Master Quorum Isolation & Split-Brain Prevention | Provisions dedicated 3-pod master-eligible NodeSets with strict pod anti-affinity, cluster.initial_master_nodes bootstrapping, and no data roles to prevent quorum split-brain. | Co-locates master and data roles on the same worker pods; heavy search queries exhaust CPU and memory, dropping master heartbeats and triggering false leader elections. | Configures 2-pod or even-numbered master nodes without odd-quorum voting configuration, inducing split-brain partitions during network blips. | Deploys a single master node pod without quorum redundancy; any worker node rescheduling halts cluster state management cluster-wide. |
| Persistent Volume Claim (PVC) & Local NVMe Storage Classes | Engineers low-latency NVMe storage classes via DirectPV/local-volume CSI with dynamic volume expansion, volumeClaimTemplates, and S3 searchable snapshot tiering. | Relies on default network-attached block storage (e.g., standard gp3/pd-standard) with unoptimized IOPS, causing disk queue latency spikes during bulk indexing. | Mounts shared NFS or cluster-wide network filesystems across Elasticsearch pods, causing Lucene index lock contentions and segment corruption. | Uses ephemeral hostPath or emptyDir storage without persistent volume durability; worker node restarts permanently wipe all indexed documents. |
| Pod Anti-Affinity & Multi-Zone Shard Allocation Awareness | Enforces topologySpreadConstraints and hard podAntiAffinity across availability zones, synchronizing Kubernetes topology labels with cluster.routing.allocation.awareness.attributes: zone. | Deploys pods without zone awareness; primary and replica shards land on worker nodes in the same cloud availability zone, creating single-point-of-failure exposure. | Relies on basic Kubernetes default scheduling; multiple master and data pods schedule onto the same physical hypervisor node. | Ignores affinity and multi-zone placement entirely; entire clusters go offline when Kubernetes reschedules pods to a single drained node. |
| Rolling Zero-Downtime Stateful Upgrades & PDBs | Implements strict PodDisruptionBudgets (maxUnavailable=1 for masters and data tiers), automated shard reallocation disabling (cluster.routing.allocation.enable: "primaries"), and preStop sync hooks. | Executes blind kubectl rollout restart, restarting multiple data pods simultaneously, triggering massive cluster-wide shard recovery storms and degraded query latency. | Schedules multi-hour maintenance downtime to stop and recreate StatefulSets sequentially during off-peak hours. | Updates container image tags directly in production manifests without PDBs; Kubernetes terminates multiple pods concurrently, causing red cluster state. |
| Native Cloud Observability & Metricbeat Exporters | Integrates Elastic Agent/Metricbeat DaemonSets, Prometheus Elasticsearch Exporter, and curated Grafana dashboards monitoring JVM GC pauses, Lucene segment merges, and indexing queue depths. | Relies only on basic Kubernetes pod CPU and memory utilization metrics, remaining blind to Lucene segment counts, circuit breaker trips, and unassigned shards. | Injects generic APM agents into Elasticsearch containers, consuming critical JVM heap without tracking cluster state publication latencies. | No monitoring configured; discovers cluster degradation only after customer support tickets report search timeouts and 503 errors. |
Kubernetes Failure Modes
Critical Elasticsearch on Kubernetes Risks We Eliminate
Running distributed Lucene clusters inside containerized Kubernetes environments introduces quorum split-brain, PVC volume attachment, and disk watermark failure vectors. We engineer resilience into every layer to eliminate these production risks:
Master Node Pod Preemption Causing Quorum Split-Brain
During Kubernetes cluster autoscaling or node pool upgrades, master-eligible Elasticsearch pods are evicted concurrently without honoring quorum preservation. The master quorum drops below 2 out of 3, freezing cluster state publication and halting all cluster-wide write and indexing operations.
JusDB provisions dedicated master node pools with PodDisruptionBudgets (maxUnavailable=1), isolates masters from data workloads, and configures cluster.initial_master_nodes and voting configuration exclusions before node maintenance.
Disk Watermark Breach Freezing Kubernetes Data Pods
When data pods breach high disk watermarks (85% low, 90% high, 95% flood stage) due to high-volume indexing, Elasticsearch automatically switches indices into read-only mode (index.blocks.read_only_allow_delete). Kubernetes PVC expansion cannot occur without restarting pods unless CSI volume expansion is properly configured.
JusDB deploys automated dynamic PVC storage expansion via Kubernetes CSI drivers, tunes low/high/flood watermarks, and establishes automated ILM rollover policies moving older segments to S3 searchable snapshot tiers.
Cascading Shard Relocation During Uncoordinated Node Draining
When Kubernetes worker nodes are drained without disabling shard allocation, Elasticsearch immediately detects missing data pods and initiates aggressive cluster-wide replica rebalancing. Massive cross-node network traffic saturates cluster interfaces, creating query latency spikes and node heartbeat timeouts.
JusDB implements preStop lifecycle hooks and automated drain runbooks that set cluster.routing.allocation.enable: 'primaries' prior to pod termination, preventing catastrophic cascading shard movements during routine node cycling.
Our DBREs execute non-blocking diagnostic commands across Kubernetes CRD controllers and internal Elasticsearch cluster settings to audit node set health and zone allocation awareness:
Audits ECK custom resource health, node set reconciliation phases, pod scheduling states, and worker node distributions across Kubernetes.
# Audit ECK Elasticsearch custom resource reconciliation and wide pod topology kubectl get elasticsearch,pods -l common.k8s.elastic.co/type=elasticsearch -o wide
Verifies that cluster shard routing allocation awareness attributes correctly mirror underlying Kubernetes availability zones.
# Verify Kubernetes availability zone shard allocation awareness configuration curl -s -X GET "http://localhost:9200/_cluster/settings?include_defaults=true" | jq '.defaults.cluster.routing.allocation.awareness'
FAQ
Elasticsearch on Kubernetes — Frequently Asked Questions
Common questions about deploying and operating Elasticsearch on Kubernetes with ECK.
Ready to Run Elasticsearch on Kubernetes?
Let our experts deploy and manage production-grade Elasticsearch clusters on Kubernetes with ECK, auto-scaling, and zero-downtime rolling upgrades.
Related Elasticsearch Services
Explore more ways our Elasticsearch experts can help with your database infrastructure.