Considering Apache Druid?
- ▸ Time-series OLAP at scale — observability or ad-tech workload with billions of events / day and the roll-up storage savings are the reason you're looking past ClickHouse and Pinot.
- ▸ Kafka indexing topology — supervisor tasks aren't auto-balancing the way you expected and segment compaction is becoming the operational bottleneck.
- ▸ Imply Polaris vs self-managed — the TCO model needs real numbers, and the team is debating whether the operational savings justify the managed-service premium.
JusDB Apache Druid specialists design, deploy, and operate real-time OLAP at scale. See Druid consulting →
Apache Druid, time-series OLAP at any scale.
Apache Druid is an open-source, distributed real-time OLAP datastore engineered for sub-second analytical queries on high-throughput event data. It decouples ingestion (Overlord + MiddleManagers), storage (Historical nodes + deep storage), and query processing (Brokers + Routers), leveraging ingestion-time roll-up pre-aggregation and time-partitioned columnar segments. JusDB provides 24/7 Druid DBRE: supervisor tuning, auto-compaction, query optimization, and guaranteed <15m P1 incident response.
Kafka indexing supervisors, roll-up pre-aggregation, time-partitioned segments, and the Coordinator + Overlord + Historical + Broker + Router topology — purpose-built for time-series OLAP at billions-of-events scale.
Apache Druid · historical+realtime
Time-partitioned segments · roll-up
0.00k
15ms
0k
0.1M
Event Ingestion
0.00M events/s[OK] segment: clicks_2026-06-19 handed off to historical
[INF] realtime: kafka indexing task consuming, lag 1s
[OK] compaction: merged 1,840 small segments → 12
[INF] coordinator: segment balance even across tiers
Representative fleet view · illustrative metrics
0M+
Events / sec Ingested
0.99%
Uptime SLA
0×
Median Query Speedup
0×
Roll-up Storage Win
Apache Druid service paths
Druid Consulting
Kafka indexing topology, roll-up strategy, segment granularity, Coordinator / Historical tier sizing, Imply Polaris vs self-managed economics — written advisory deliverables.
Druid vs Pinot
Side-by-side comparison — roll-up vs star-tree, Kafka indexing vs LLC, multi-tenancy models, Imply Polaris vs StarTree Cloud, when each one wins.
What we do
What we build with Apache Druid
From cluster design to production query tuning — end-to-end Druid expertise.
Real-Time Kafka Ingestion
Kafka indexing service auto-scales supervisor tasks across MiddleManagers; second-level freshness from topic to query-ready segments.
Roll-Up Pre-Aggregation
Destructive aggregation at ingestion time — 10-100x storage reduction for time-series workloads where raw rows aren't needed downstream.
Time-Partitioned Segments
Segments natively partitioned by time interval — query pruning, retention policies, and compaction all operate on time-aligned units.
Tier-Decoupled Architecture
Coordinator + Overlord + Historical + Broker + Router tiers scale independently — match infrastructure to actual workload shape.
Deep Storage + Hot Tiers
Deep storage on S3/HDFS/GCS plus hot Historical-node caching — predictable retention with cost-aware tiering.
Imply Polaris Operations
Managed-Druid SaaS with Pivot visualisation included — fast time-to-value when operational burden is the dominant cost.
Time-series OLAP performance
Druid expertise
for event-scale analytics
We tune roll-up dimensions and segment granularity, balance Kafka indexing supervisors, and right-size the Historical and Broker tiers so time-series queries return in sub-second time even as event volume grows into the billions.
Time-Series Performance
After tuning40×
Median speedup
20×
Roll-up storage win
Real cases
Queries we've transformed
6,000ms
90ms
Raw events stored — billions of un-aggregated rows
The fix
Enabled roll-up at ingest — 20× less storage, faster scans
4,800ms
70ms
1,800 tiny segments — broker fan-out overhead
The fix
Configured auto-compaction to merge into optimal segments
1,900ms
35ms
Repeated dashboard queries recomputed every time
The fix
Enabled broker result cache for hot time-series queries
0.00%
Cluster Uptime
<0s
Failover RTO
0s
Ingestion Lag
High availability
Always on. Tier-decoupled.
Druid's tiers fail over independently — Coordinator and Overlord run in active/standby pairs, Historicals serve replicated segments, and deep storage means any lost node is re-served without data loss.
Incident response
A supervisor-stall P1, handled in under 15 minutes.
When a Kafka indexing supervisor stalls and ingestion lag spikes, a named Druid engineer responds — not a ticket queue. We reset the supervisor, rebalance tasks, and clear the backlog online, with a blameless postmortem after.
Time-series dashboard p99 > 6s — analysts blocked
Named OLAP engineer in under 15 min, not a ticket queue
Raw events stored — no roll-up, 1,800 tiny segments
Enabled roll-up at ingest + auto-compaction policy
Storage 20× smaller, p99 6s → 90ms — total 14 min
Pre-Migration Assessment
Data warehouse / Pinot → Apache Druid
Estimated cutover window: < 10 minutes
Migration
Move to Apache Druid without the downtime
Pinot, ClickHouse or a homegrown time-series store → Druid. We design roll-up and segment strategy, backfill historical segments in parallel, stand up Kafka indexing supervisors, and cut over once query results reconcile.
Real-Time OLAP Architecture
Apache Druid Production Failure Modes
Distributed OLAP architectures face distinct supervisor worker exhaustion, segment explosion, and coordinator locking pressures. Here is how JusDB DBREs diagnose and eliminate Druid's most critical production failure modes.
Kafka Indexing Supervisor Task Stalls & Ingestion Lag Cascades
When Kafka topic partitions experience sudden burst traffic or malformed event payloads, Druid indexing tasks fail or exceed taskDuration limits without committing offsets. The Overlord repeatedly spawns failing replacement tasks, exhausting middleManager worker slots and accumulating massive ingestion lag that blinds downstream analytical dashboards.
JusDB DBREs tune supervisor taskCount and taskDuration thresholds, configure automated dead-letter routing for malformed payloads, implement dynamic middleManager auto-scaling on Kubernetes, and set up Prometheus alert rules for early consumer lag detection.
Small Segment Explosion & Historical Node Heap Memory Exhaustion
Frequent real-time segment generation without timely auto-compaction produces millions of tiny (5MB-20MB) segments. Each segment requires metadata storage in ZooKeeper, the Coordinator, and Historical node JVM heap memory, causing garbage collection pauses to spike into tens of seconds and paralyzing query broker scatter-gather stages.
We implement aggressive Druid auto-compaction policies targeting the 400MB-600MB optimal segment size, configure intelligent tier-based segment caching rules, and tune G1GC parameters across Historical and Broker nodes to eliminate pause degradation.
Coordinator & Overlord Leader Election Deadlocks under Metadata Pressure
Under intense concurrent segment handoff from middleManagers to deep storage, the metadata store (PostgreSQL/MySQL) and ZooKeeper experience severe lock contention. The Coordinator loses its ephemeral ZooKeeper leader lock, triggering repeated leader election cycles and halting all segment assignment and historical balancing.
JusDB isolates metadata store connections with PgBouncer connection multiplexing, scales ZooKeeper quorums with dedicated NVMe storage, and configures Coordinator balancing intervals with jitter to prevent thunderous herd rebalance storms.
Cluster Telemetry
Production Apache Druid Diagnostic Runbooks
Non-blocking Overlord API queries and Druid SQL system catalog diagnostics executed by JusDB DBREs during incident triage to isolate Kafka ingestion lag, supervisor stalls, and Historical segment heap pressure.
Queries Kafka supervisor consumer lag across topic partitions and verifies active worker slot allocations on middleManagers.
# Query active supervisor status, consumer lag, and partition assignments
curl -s http://localhost:8888/druid/indexer/v1/supervisor/telemetry_events_kafka/status | jq '{
id: .id,
generationTime: .generationTime,
state: .payload.state,
aggregateLag: .payload.aggregateLag,
partitionLag: .payload.partitionLag
}'
# Inspect MiddleManager active tasks, capacity, and worker slots
curl -s http://localhost:8888/druid/indexer/v1/workers | jq '.[] | {
host: .worker.host,
capacity: .worker.capacity,
currCapacityUsed: .currCapacityUsed
}'Surfaces fragmented small segment counts by datasource and audits Historical node memory utilization.
-- Query segment count, total size, and average segment size per datasource
SELECT datasource,
COUNT(*) AS total_segments,
ROUND(SUM("size") / 1024 / 1024, 2) AS total_size_mb,
ROUND(AVG("size") / 1024 / 1024, 2) AS avg_segment_size_mb
FROM sys.segments
WHERE is_active = 1
GROUP BY datasource
ORDER BY total_segments DESC;
-- Identify Historical nodes approaching max heap memory usage
SELECT host, curr_size, max_size,
ROUND((curr_size * 100.0) / max_size, 2) AS utilization_pct
FROM sys.servers
WHERE server_type = 'historical';Comparative Analysis
Apache Druid DBRE: Evaluation Matrix
How JusDB specialized Apache Druid reliability engineering compares against Imply Polaris default managed tier and in-house generalists.
| Evaluation Vector | JusDB Druid DBRE | Imply Polaris / Default | In-House Generalists |
|---|---|---|---|
| Kafka Ingestion & Supervisor Auto-Scaling | Automated supervisor task balancing, dynamic task autoscaling aligned with Kafka partitions, and zero-data-loss checkpointing | Basic console supervisor creation; dynamic task scaling and partition rebalancing require manual administrative intervention | Under-provisioned middleManager slots causing unassigned tasks, consumer group offset drift, and hours of real-time ingestion lag |
| Ingestion Roll-Up & Storage Footprint | Aggressive destructive roll-up modeling and custom metric aggregations delivering 10x-100x storage savings for time-series events | Default ingestion schemas retain high-cardinality raw dimensions, driving up cloud storage and Historical node memory requirements | Accidentally ingesting unique timestamps or request UUIDs as dimensions, completely defeating segment roll-up pre-aggregation |
| Time-Partitioned Segments & Compaction | Automated auto-compaction rules, segment size calibration (400-600MB sweet spot), and intelligent multi-interval retention policies | Manual compaction setup; fragmented small segments degrade query pruning and exhaust Historical heap memory | Segment explosion with millions of tiny 5MB segments choking the Coordinator coordinator and Broker scatter-gather phases |
| Tier-Decoupled Topology & Deep Storage | Independent scaling of Coordinator, Overlord, Historical, Broker, and Router tiers with tiered cold/hot S3/GCS deep-storage caching | Standard bundled compute tiers where query and ingestion resources cannot be decoupled for asymmetric workloads | Co-locating Historical and MiddleManager services on identical compute instances, resulting in severe CPU starvation during bulk loads |
| SQL Engine & Query Vectorization | Cost-based query plan analysis, Calcite SQL optimization, vectorized expression tuning, and JDBC connection pool multiplexing for BI tools | Standard SQL-to-native compilation without subquery rewrite optimization or custom query cache tuning | Heavy unindexed multi-join SQL queries submitted by BI tools triggering full Historical scans and Broker memory out-of-memory crashes |
| 24/7 Production DBRE & Sub-15m P1 SLA | Senior Apache Druid DBREs on-call 24/7/365 with contractual <15m P1 incident response and zero-ticket escalation queues | Standard enterprise cloud ticketing with 1-to-2 hour initial response windows on high-severity incidents | Platform teams manually debugging Coordinator thread dumps and ZooKeeper connection loss at 3 AM with production dashboards down |
FAQ
Apache Druid — common questions
Get started
Ready to evaluate Druid?
Book a 30-minute scoping call. We'll review your workload shape, the roll-up strategy, and the managed-vs-self-managed decision before any statement of work.
Explore Our Apache Druid Services
Explore more ways our Apache Druid experts can help with your database infrastructure.