Sound familiar?
- ▸ Roll-up granularity decisions are blocking the schema design — the team needs guidance on which dimensions to keep raw and which to aggregate, and the trade-off is hard to articulate without query-pattern data.
- ▸ Kafka indexing topology isn't scaling cleanly — supervisor tasks lag behind partition growth, segment-fetch costs dominate, and the team needs a defensible architecture call.
- ▸ Imply Polaris vs self-managed — finance wants a defensible TCO model before signing the renewal, and the workload-shape assumptions need to be tested against real data.
JusDB Apache Druid consultants give you the written decision document — not a Slack-thread opinion. Book a Druid architecture review →
Strategic advisory — not execution
Apache Druid Consulting Services
In short: Apache Druid consulting is strategic advisory delivered as written recommendations — roll-up dimension selection, Kafka indexing supervisor topology, Coordinator/Historical/Broker tier sizing, segment compaction strategy, deep-storage choice, and Imply Polaris vs self-managed economics. You need it when roll-up granularity, Kafka indexing scale, or the Polaris-vs-self-managed TCO call is blocking your design.
Roll-up dimension selection, Kafka indexing supervisor topology, Coordinator / Historical tier sizing, Imply Polaris vs self-managed economics, and segment compaction strategy. See the Druid hub for the broader services overview, or the Druid vs Pinot comparison for the side-by-side decision matrix.
JusDB provides enterprise Apache Druid consulting to architect high-concurrency streaming OLAP systems, optimize segment sizing (5M rows / 500MB target), and design ingestion-time roll-up schemas achieving 10x to 100x storage efficiency. Certified Database Reliability Engineers calibrate MiddleManager Kafka supervisors, tune Historical query caches, and automate deep storage tiering, backed by contractual 15-minute emergency SLAs.
Advisory scope
What our Druid consulting covers
Each deliverable is a written decision document, sized topology proposal, or costed trade-off analysis.
Roll-Up Dimension Selection
Which dimensions stay raw, which get rolled up, and at what granularity — driven by audited query patterns, not blanket aggregation.
Kafka Indexing Topology
Supervisor task count aligned to Kafka partitions, MiddleManager / Indexer capacity sizing, late-data handling, replay strategy.
Tier Sizing
Coordinator + Overlord + Historical + Broker + Router sizing based on query QPS, scan size, and ingestion throughput.
Imply Polaris vs Self-Managed
TCO modelling against actual workload — Polaris managed-SaaS economics vs self-managed-on-K8s reserved-instance pricing.
Segment Compaction Strategy
Compaction rules per datasource — target segment size, time interval, granularity — designed against query latency budget.
Engine Decision Matrix
Druid vs Pinot vs ClickHouse vs StarRocks for the specific workload — modelled against latency, concurrency, and retention requirements.
Greenfield & Migration Strategy (planning only)
Greenfield architecture or migration planning from an existing OLAP engine — topology spec, capacity model, security baseline, ops runbook. Execution is a separate engagement.
How JusDB Apache Druid Consulting compares to alternative models.
Standard cloud hosting support and generic IT contractors lack deep Apache Druid internals, segment sizing mechanics (5M rows / 500MB target), rollup pre-aggregation strategies, and continuous DBRE reliability ownership. Here is how our certified Druid specialists compare:
| Evaluation Vector | JusDB DBRE | In-House DBA | Legacy Agency | Developer Generalist |
|---|---|---|---|---|
| Segment Sizing & Sharding Architecture (5M rows / 500MB) | Calibrates segment partitioning targeting the golden 5M rows / 500MB standard, configures dynamic partition pruning, and sets shardSpecs to maximize Historical scan vectorization without memory bloat. | Permits unconstrained segment generation with tiny 10MB segments or massive 5GB segments, overloading Coordinator metadata heaps and stalling Broker scatter-gather query fans. | Treats Druid like a traditional relational database; ignores segment granularity and interval boundaries, leading to segment scatter across hundreds of Historical nodes. | Leaves ingestion specs at default partition settings; creates millions of sub-megabyte segments that crash Coordinator loops and exhaust ZooKeeper state nodes. |
| Ingestion-Time Rollup & Pre-Aggregation Strategy | Architects destructive ingestion-time rollup with calibrated timestamp truncation and metric sketches (HyperLogLog, Theta sketches), achieving 10x-100x storage reduction and sub-second aggregations. | Ingests raw millisecond timestamps and high-cardinality transaction IDs as dimensions, completely defeating segment rollup and exploding cloud deep storage costs. | Disables rollup out of caution, forcing Historical nodes to scan billions of raw event rows for simple dashboard sum and count queries. | Applies rollup blindly without understanding downstream drill-down requirements, irreversibly destroying raw dimensions needed for business reporting. |
| Tiered Historical & Broker Decoupled Topology | Architects multi-tier Historical topologies (hot NVMe vs cold SSD/object store) with automated drop/load retention rules, query routing isolation, and independent Broker query pooling. | Runs a single monolithic Historical tier, storing multi-year cold archives on expensive hot compute and risking memory starvation during seasonal reporting peaks. | Co-locates Brokers and Historicals on identical instances, causing query planning thread contention and uncoordinated segment cache eviction. | Over-provisions Broker memory while starving Historical processing thread buffers, resulting in severe GroupBy query timeouts. |
| MiddleManager Kafka/Kinesis Supervisor Concurrency | Calibrates supervisor task counts, worker task slots, and heap/off-heap direct buffers aligned with streaming partitions, ensuring zero-lag ingestion and seamless segment handoffs. | Overallocates supervisor tasks beyond physical MiddleManager worker slots, leaving real-time tasks queued in pending state and falling behind Kafka retention. | Implements custom cron-based micro-batch scripts over HTTP instead of native streaming supervisors, causing frequent data loss and ingestion stalls. | Assigns arbitrary maxRowsInMemory thresholds, causing premature intermediate segment spilling to disk and frequent OutOfMemoryErrors on indexing workers. |
| Cloud Deep Storage Lifecycle (S3/GCS/Azure Blob) | Configures reliable cloud deep storage integrations (S3/GCS/Azure Blob) with lifecycle archive tiers, atomic segment push/pull mechanics, and resilient Coordinator kill-task cleanup policies. | Fails to automate segment deletion or kill-tasks in deep storage, leaving orphan segment files accumulating hundreds of terabytes in cloud object storage. | Mounts shared NFS or local POSIX storage as deep storage, creating single-point-of-failure storage bottlenecks and cluster-wide crash risks. | Deploys Druid without durable deep storage configuration, risking catastrophic cluster-wide data loss upon worker node restarts. |
| Enterprise Security, TLS & Kerberos/RBAC | Implements end-to-end mTLS encryption across all Druid tiers, Kerberos/OIDC authentication, role-based access control (RBAC), and SQL query-level column masking. | Configures basic plaintext HTTP internal communication, leaving cluster control APIs and inter-node segment replication exposed on local networks. | Relies on external edge reverse proxies alone without configuring internal Druid authenticator extensions or role-based datasource permissions. | Disables authentication extensions entirely to bypass configuration complexity, exposing the Druid Coordinator and Router consoles to unauthorized access. |
Apache Druid Engine Failure Modes
Critical Apache Druid Outage Modes We Eliminate
High-throughput streaming Apache Druid clusters encounter severe availability and latency risks when small segments explode Historical server memory, Coordinator locks freeze segment handoffs, or MiddleManager workers exhaust JVM heaps. Our DBREs resolve these breakdown modes:
Small Segment Explosion Saturating Historical Server Memory
Ingestion pipelines generating millions of uncompacted sub-100MB segments force Historical nodes to maintain massive segment metadata trees in heap. Heap exhaustion triggers prolonged stop-the-world GC pauses, Coordinator assignment lags, and cluster-wide Broker query timeouts.
JusDB DBREs calibrate automated Coordinator compaction rules to consolidate small fragments into 400-600MB segments, optimize partition sharding, and tune Historical off-heap direct memory buffers for optimal scan vectorization.
Coordinator Metadata Lock Contention Freezing Segment Handoffs
Aggressive metadata polling combined with heavy real-time publishing tasks causes table-level lock contention on PostgreSQL/MySQL metadata stores. Coordinator fails to acknowledge segment handoffs from MiddleManagers, causing streaming tasks to stall and memory buffers to overflow.
JusDB optimizes metadata database connection pooling, adjusts druid.coordinator.period and handoff poll intervals, and decouples segment management loops to ensure non-blocking segment handoffs.
MiddleManager Worker Heap Exhaustion Crashing Indexing Tasks
High-cardinality dimension bursts or improperly configured maxRowsInMemory force MiddleManager worker peons to exceed physical JVM heap ceilings during ingestion rollup. Tasks crash with OutOfMemoryError, forcing supervisor re-tries and driving up Kafka consumer lag.
JusDB DBREs tune druid.indexer.fork.property JVM arguments, calibrate maxBytesInMemory thresholds, and configure proactive worker heap monitoring to guarantee smooth indexing handoffs.
Our Apache Druid DBREs execute non-blocking telemetry inspections to verify segment size distributions, compaction health, and streaming supervisor states without interrupting analytical query workloads:
Queries the Druid system catalog to evaluate segment count, average segment size in megabytes, and average row count grouped by datasource to detect small-segment fragmentation.
curl -s -X POST "http://localhost:8888/druid/v2/sql" \
-H "Content-Type: application/json" \
-d '{"query": "SELECT datasource, COUNT(*) AS num_segments, AVG(\"size\")/1048576 AS avg_size_mb, AVG(num_rows) AS avg_rows FROM sys.segments WHERE is_active = 1 GROUP BY 1"}'Inspects active Kafka and Kinesis indexing supervisors, reporting status states, task health, generation IDs, and consumer lag offsets across streaming pipelines.
curl -s -X GET "http://localhost:8888/druid/indexer/v1/supervisor"
FAQ
Druid consulting — common questions
Ready to make the call on Druid?
Book a 30-minute scoping call. We'll tell you which engagement shape fits and what the deliverable will look like — before any statement of work.
Related Apache Druid Services
Explore more ways our Apache Druid experts can help with your database infrastructure.