Free audit

View Audit Scope

Considering Apache Druid?

  • ▸ Time-series OLAP at scale — observability or ad-tech workload with billions of events / day and the roll-up storage savings are the reason you're looking past ClickHouse and Pinot.
  • ▸ Kafka indexing topology — supervisor tasks aren't auto-balancing the way you expected and segment compaction is becoming the operational bottleneck.
  • ▸ Imply Polaris vs self-managed — the TCO model needs real numbers, and the team is debating whether the operational savings justify the managed-service premium.

JusDB Apache Druid specialists design, deploy, and operate real-time OLAP at scale. See Druid consulting →

Druid · Kafka Indexing · Roll-Up
Real-Time OLAP, Time-Series-First

Apache Druid, time-series OLAP at any scale.

Executive Direct Answer · Apache Druid Architecture

Apache Druid is an open-source, distributed real-time OLAP datastore engineered for sub-second analytical queries on high-throughput event data. It decouples ingestion (Overlord + MiddleManagers), storage (Historical nodes + deep storage), and query processing (Brokers + Routers), leveraging ingestion-time roll-up pre-aggregation and time-partitioned columnar segments. JusDB provides 24/7 Druid DBRE: supervisor tuning, auto-compaction, query optimization, and guaranteed <15m P1 incident response.

Architecture: Decoupled Multi-Tier OLAP·Ingestion: Native Kafka/Kinesis Supervisors·Storage: Deep Storage (S3/GCS) + Historical Caching·Compression: 10x-100x Roll-Up Pre-Aggregation·P1 SLA: <15 Min

Kafka indexing supervisors, roll-up pre-aggregation, time-partitioned segments, and the Coordinator + Overlord + Historical + Broker + Router topology — purpose-built for time-series OLAP at billions-of-events scale.

JUSDB_DRUID_PROD
LIVE

Apache Druid · historical+realtime

Time-partitioned segments · roll-up

Tuned
Queries / sec

0.00k

Query p99

15ms

Segments

0k

Ingest events / sec

0.1M

Event Ingestion

0.00M events/s

[OK] segment: clicks_2026-06-19 handed off to historical

[INF] realtime: kafka indexing task consuming, lag 1s

[OK] compaction: merged 1,840 small segments → 12

[INF] coordinator: segment balance even across tiers

Representative fleet view · illustrative metrics

0M+

Events / sec Ingested

0.99%

Uptime SLA

0×

Median Query Speedup

0×

Roll-up Storage Win

What we do

What we build with Apache Druid

From cluster design to production query tuning — end-to-end Druid expertise.

Real-Time Kafka Ingestion

Kafka indexing service auto-scales supervisor tasks across MiddleManagers; second-level freshness from topic to query-ready segments.

Roll-Up Pre-Aggregation

Destructive aggregation at ingestion time — 10-100x storage reduction for time-series workloads where raw rows aren't needed downstream.

Time-Partitioned Segments

Segments natively partitioned by time interval — query pruning, retention policies, and compaction all operate on time-aligned units.

Tier-Decoupled Architecture

Coordinator + Overlord + Historical + Broker + Router tiers scale independently — match infrastructure to actual workload shape.

Deep Storage + Hot Tiers

Deep storage on S3/HDFS/GCS plus hot Historical-node caching — predictable retention with cost-aware tiering.

Imply Polaris Operations

Managed-Druid SaaS with Pivot visualisation included — fast time-to-value when operational burden is the dominant cost.

Time-series OLAP performance

Druid expertise
for event-scale analytics

We tune roll-up dimensions and segment granularity, balance Kafka indexing supervisors, and right-size the Historical and Broker tiers so time-series queries return in sub-second time even as event volume grows into the billions.

Roll-up dimension and metric design at ingestion time
Segment granularity and time-partition tuning
Kafka indexing supervisor balancing and auto-scaling
Compaction strategy for long-retention segments
Coordinator / Historical / Broker tier right-sizing
Prometheus + Grafana query-latency and segment dashboards

Time-Series Performance

After tuning
Roll-up pre-aggregation coverage0%
Segment compaction efficiency0%
Broker result-cache hit rate0%
Tiered storage data locality0%

40×

Median speedup

20×

Roll-up storage win

Real cases

Queries we've transformed

No Roll-Up

6,000ms

90ms

Raw events stored — billions of un-aggregated rows

The fix

Enabled roll-up at ingest — 20× less storage, faster scans

Too Many Small Segments

4,800ms

70ms

1,800 tiny segments — broker fan-out overhead

The fix

Configured auto-compaction to merge into optimal segments

Broker Not Caching

1,900ms

35ms

Repeated dashboard queries recomputed every time

The fix

Enabled broker result cache for hot time-series queries

Cluster ACTIVEMaster + Query + Data tiers

0.00%

Cluster Uptime

<0s

Failover RTO

0s

Ingestion Lag

coordinator / overlord
MASTERONLINE
broker-01..02 · 8082
QUERYONLINE
historical + middlemgr
DATAONLINE

High availability

Always on. Tier-decoupled.

Druid's tiers fail over independently — Coordinator and Overlord run in active/standby pairs, Historicals serve replicated segments, and deep storage means any lost node is re-served without data loss.

Replicated segments across Historical nodes
Active/standby Coordinator & Overlord with leader election
Deep-storage durability on S3 / GCS / HDFS
Independent tier scaling — match infra to workload
Tested recovery runbooks for indexing & segment failures

Incident response

A supervisor-stall P1, handled in under 15 minutes.

When a Kafka indexing supervisor stalls and ingestion lag spikes, a named Druid engineer responds — not a ticket queue. We reset the supervisor, rebalance tasks, and clear the backlog online, with a blameless postmortem after.

P1 alert → named Druid engineer paged in under 15 minutes
Root cause via Overlord logs, supervisor status & Grafana
Supervisor reset + task rebalance — no query downtime
Blameless postmortem with a prevention plan
Live incident replayP1 → resolved · ~14 min
1
00:00Alert fired

Time-series dashboard p99 > 6s — analysts blocked

2
00:03On-call paged

Named OLAP engineer in under 15 min, not a ticket queue

3
00:07Root cause

Raw events stored — no roll-up, 1,800 tiny segments

4
00:11Fix applied

Enabled roll-up at ingest + auto-compaction policy

5
00:14Resolved

Storage 20× smaller, p99 6s → 90ms — total 14 min

Pre-Migration Assessment

Data warehouse / Pinot → Apache Druid

READY
Datasource & roll-up spec design0%
Batch backfill (deep storage)0%
Realtime Kafka ingestion catch-up0%
Cutover readiness0%

Estimated cutover window: < 10 minutes

Migration

Move to Apache Druid without the downtime

Pinot, ClickHouse or a homegrown time-series store → Druid. We design roll-up and segment strategy, backfill historical segments in parallel, stand up Kafka indexing supervisors, and cut over once query results reconcile.

Roll-up & segment-granularity modeling for your workload
Parallel historical backfill plus Kafka indexing sync
Query-result reconciliation before cutover
Self-managed (K8s/EC2) & Imply Polaris targets
Plan My Migration

Real-Time OLAP Architecture

Apache Druid Production Failure Modes

Distributed OLAP architectures face distinct supervisor worker exhaustion, segment explosion, and coordinator locking pressures. Here is how JusDB DBREs diagnose and eliminate Druid's most critical production failure modes.

High Severity (P1)

Kafka Indexing Supervisor Task Stalls & Ingestion Lag Cascades

When Kafka topic partitions experience sudden burst traffic or malformed event payloads, Druid indexing tasks fail or exceed taskDuration limits without committing offsets. The Overlord repeatedly spawns failing replacement tasks, exhausting middleManager worker slots and accumulating massive ingestion lag that blinds downstream analytical dashboards.

JusDB Engineering Mitigation

JusDB DBREs tune supervisor taskCount and taskDuration thresholds, configure automated dead-letter routing for malformed payloads, implement dynamic middleManager auto-scaling on Kubernetes, and set up Prometheus alert rules for early consumer lag detection.

High Severity (P1)

Small Segment Explosion & Historical Node Heap Memory Exhaustion

Frequent real-time segment generation without timely auto-compaction produces millions of tiny (5MB-20MB) segments. Each segment requires metadata storage in ZooKeeper, the Coordinator, and Historical node JVM heap memory, causing garbage collection pauses to spike into tens of seconds and paralyzing query broker scatter-gather stages.

JusDB Engineering Mitigation

We implement aggressive Druid auto-compaction policies targeting the 400MB-600MB optimal segment size, configure intelligent tier-based segment caching rules, and tune G1GC parameters across Historical and Broker nodes to eliminate pause degradation.

Critical (P1)

Coordinator & Overlord Leader Election Deadlocks under Metadata Pressure

Under intense concurrent segment handoff from middleManagers to deep storage, the metadata store (PostgreSQL/MySQL) and ZooKeeper experience severe lock contention. The Coordinator loses its ephemeral ZooKeeper leader lock, triggering repeated leader election cycles and halting all segment assignment and historical balancing.

JusDB Engineering Mitigation

JusDB isolates metadata store connections with PgBouncer connection multiplexing, scales ZooKeeper quorums with dedicated NVMe storage, and configures Coordinator balancing intervals with jitter to prevent thunderous herd rebalance storms.

Cluster Telemetry

Production Apache Druid Diagnostic Runbooks

Non-blocking Overlord API queries and Druid SQL system catalog diagnostics executed by JusDB DBREs during incident triage to isolate Kafka ingestion lag, supervisor stalls, and Historical segment heap pressure.

Kafka Supervisor Lag & Worker Slots Telemetry
Overlord API · Real-time

Queries Kafka supervisor consumer lag across topic partitions and verifies active worker slot allocations on middleManagers.

# Query active supervisor status, consumer lag, and partition assignments
curl -s http://localhost:8888/druid/indexer/v1/supervisor/telemetry_events_kafka/status | jq '{
  id: .id,
  generationTime: .generationTime,
  state: .payload.state,
  aggregateLag: .payload.aggregateLag,
  partitionLag: .payload.partitionLag
}'

# Inspect MiddleManager active tasks, capacity, and worker slots
curl -s http://localhost:8888/druid/indexer/v1/workers | jq '.[] | {
  host: .worker.host,
  capacity: .worker.capacity,
  currCapacityUsed: .currCapacityUsed
}'
Segment Compaction & Historical Memory Diagnostics
sys.segments · Zero-overhead

Surfaces fragmented small segment counts by datasource and audits Historical node memory utilization.

-- Query segment count, total size, and average segment size per datasource
SELECT datasource,
       COUNT(*) AS total_segments,
       ROUND(SUM("size") / 1024 / 1024, 2) AS total_size_mb,
       ROUND(AVG("size") / 1024 / 1024, 2) AS avg_segment_size_mb
FROM sys.segments
WHERE is_active = 1
GROUP BY datasource
ORDER BY total_segments DESC;

-- Identify Historical nodes approaching max heap memory usage
SELECT host, curr_size, max_size,
       ROUND((curr_size * 100.0) / max_size, 2) AS utilization_pct
FROM sys.servers
WHERE server_type = 'historical';

Comparative Analysis

Apache Druid DBRE: Evaluation Matrix

How JusDB specialized Apache Druid reliability engineering compares against Imply Polaris default managed tier and in-house generalists.

Evaluation VectorJusDB Druid DBREImply Polaris / DefaultIn-House Generalists
Kafka Ingestion & Supervisor Auto-ScalingAutomated supervisor task balancing, dynamic task autoscaling aligned with Kafka partitions, and zero-data-loss checkpointingBasic console supervisor creation; dynamic task scaling and partition rebalancing require manual administrative interventionUnder-provisioned middleManager slots causing unassigned tasks, consumer group offset drift, and hours of real-time ingestion lag
Ingestion Roll-Up & Storage FootprintAggressive destructive roll-up modeling and custom metric aggregations delivering 10x-100x storage savings for time-series eventsDefault ingestion schemas retain high-cardinality raw dimensions, driving up cloud storage and Historical node memory requirementsAccidentally ingesting unique timestamps or request UUIDs as dimensions, completely defeating segment roll-up pre-aggregation
Time-Partitioned Segments & CompactionAutomated auto-compaction rules, segment size calibration (400-600MB sweet spot), and intelligent multi-interval retention policiesManual compaction setup; fragmented small segments degrade query pruning and exhaust Historical heap memorySegment explosion with millions of tiny 5MB segments choking the Coordinator coordinator and Broker scatter-gather phases
Tier-Decoupled Topology & Deep StorageIndependent scaling of Coordinator, Overlord, Historical, Broker, and Router tiers with tiered cold/hot S3/GCS deep-storage cachingStandard bundled compute tiers where query and ingestion resources cannot be decoupled for asymmetric workloadsCo-locating Historical and MiddleManager services on identical compute instances, resulting in severe CPU starvation during bulk loads
SQL Engine & Query VectorizationCost-based query plan analysis, Calcite SQL optimization, vectorized expression tuning, and JDBC connection pool multiplexing for BI toolsStandard SQL-to-native compilation without subquery rewrite optimization or custom query cache tuningHeavy unindexed multi-join SQL queries submitted by BI tools triggering full Historical scans and Broker memory out-of-memory crashes
24/7 Production DBRE & Sub-15m P1 SLASenior Apache Druid DBREs on-call 24/7/365 with contractual <15m P1 incident response and zero-ticket escalation queuesStandard enterprise cloud ticketing with 1-to-2 hour initial response windows on high-severity incidentsPlatform teams manually debugging Coordinator thread dumps and ZooKeeper connection loss at 3 AM with production dashboards down

FAQ

Apache Druid — common questions

Get started

Ready to evaluate Druid?

Book a 30-minute scoping call. We'll review your workload shape, the roll-up strategy, and the managed-vs-self-managed decision before any statement of work.

Explore Our Apache Druid Services

Explore more ways our Apache Druid experts can help with your database infrastructure.

Compare Druid