Free audit

View Audit Scope

Sound familiar?

  • ▸ Roll-up granularity decisions are blocking the schema design — the team needs guidance on which dimensions to keep raw and which to aggregate, and the trade-off is hard to articulate without query-pattern data.
  • ▸ Kafka indexing topology isn't scaling cleanly — supervisor tasks lag behind partition growth, segment-fetch costs dominate, and the team needs a defensible architecture call.
  • ▸ Imply Polaris vs self-managed — finance wants a defensible TCO model before signing the renewal, and the workload-shape assumptions need to be tested against real data.

JusDB Apache Druid consultants give you the written decision document — not a Slack-thread opinion. Book a Druid architecture review →

Strategic advisory — not execution

Apache Druid Consulting Services

In short: Apache Druid consulting is strategic advisory delivered as written recommendations — roll-up dimension selection, Kafka indexing supervisor topology, Coordinator/Historical/Broker tier sizing, segment compaction strategy, deep-storage choice, and Imply Polaris vs self-managed economics. You need it when roll-up granularity, Kafka indexing scale, or the Polaris-vs-self-managed TCO call is blocking your design.

Roll-up dimension selection, Kafka indexing supervisor topology, Coordinator / Historical tier sizing, Imply Polaris vs self-managed economics, and segment compaction strategy. See the Druid hub for the broader services overview, or the Druid vs Pinot comparison for the side-by-side decision matrix.

Executive Direct Answer · Apache Druid Production Consulting Heuristic

JusDB provides enterprise Apache Druid consulting to architect high-concurrency streaming OLAP systems, optimize segment sizing (5M rows / 500MB target), and design ingestion-time roll-up schemas achieving 10x to 100x storage efficiency. Certified Database Reliability Engineers calibrate MiddleManager Kafka supervisors, tune Historical query caches, and automate deep storage tiering, backed by contractual 15-minute emergency SLAs.

SLA: <15-Min Sev-1·Latency: Sub-Second Analytical P99·Compression: 10x–100x Roll-up·Segment Target: 5M Rows / 500MB·Compliance: ISO 27001 & SOC 2

Advisory scope

What our Druid consulting covers

Each deliverable is a written decision document, sized topology proposal, or costed trade-off analysis.

Roll-Up Dimension Selection

Which dimensions stay raw, which get rolled up, and at what granularity — driven by audited query patterns, not blanket aggregation.

Kafka Indexing Topology

Supervisor task count aligned to Kafka partitions, MiddleManager / Indexer capacity sizing, late-data handling, replay strategy.

Tier Sizing

Coordinator + Overlord + Historical + Broker + Router sizing based on query QPS, scan size, and ingestion throughput.

Imply Polaris vs Self-Managed

TCO modelling against actual workload — Polaris managed-SaaS economics vs self-managed-on-K8s reserved-instance pricing.

Segment Compaction Strategy

Compaction rules per datasource — target segment size, time interval, granularity — designed against query latency budget.

Engine Decision Matrix

Druid vs Pinot vs ClickHouse vs StarRocks for the specific workload — modelled against latency, concurrency, and retention requirements.

Greenfield & Migration Strategy (planning only)

Greenfield architecture or migration planning from an existing OLAP engine — topology spec, capacity model, security baseline, ops runbook. Execution is a separate engagement.

Comparative Matrix · Apache Druid Consulting Architecture

How JusDB Apache Druid Consulting compares to alternative models.

Standard cloud hosting support and generic IT contractors lack deep Apache Druid internals, segment sizing mechanics (5M rows / 500MB target), rollup pre-aggregation strategies, and continuous DBRE reliability ownership. Here is how our certified Druid specialists compare:

Evaluation Vector
JusDB DBRE
In-House DBALegacy AgencyDeveloper Generalist
Segment Sizing & Sharding Architecture (5M rows / 500MB)Calibrates segment partitioning targeting the golden 5M rows / 500MB standard, configures dynamic partition pruning, and sets shardSpecs to maximize Historical scan vectorization without memory bloat.Permits unconstrained segment generation with tiny 10MB segments or massive 5GB segments, overloading Coordinator metadata heaps and stalling Broker scatter-gather query fans.Treats Druid like a traditional relational database; ignores segment granularity and interval boundaries, leading to segment scatter across hundreds of Historical nodes.Leaves ingestion specs at default partition settings; creates millions of sub-megabyte segments that crash Coordinator loops and exhaust ZooKeeper state nodes.
Ingestion-Time Rollup & Pre-Aggregation StrategyArchitects destructive ingestion-time rollup with calibrated timestamp truncation and metric sketches (HyperLogLog, Theta sketches), achieving 10x-100x storage reduction and sub-second aggregations.Ingests raw millisecond timestamps and high-cardinality transaction IDs as dimensions, completely defeating segment rollup and exploding cloud deep storage costs.Disables rollup out of caution, forcing Historical nodes to scan billions of raw event rows for simple dashboard sum and count queries.Applies rollup blindly without understanding downstream drill-down requirements, irreversibly destroying raw dimensions needed for business reporting.
Tiered Historical & Broker Decoupled TopologyArchitects multi-tier Historical topologies (hot NVMe vs cold SSD/object store) with automated drop/load retention rules, query routing isolation, and independent Broker query pooling.Runs a single monolithic Historical tier, storing multi-year cold archives on expensive hot compute and risking memory starvation during seasonal reporting peaks.Co-locates Brokers and Historicals on identical instances, causing query planning thread contention and uncoordinated segment cache eviction.Over-provisions Broker memory while starving Historical processing thread buffers, resulting in severe GroupBy query timeouts.
MiddleManager Kafka/Kinesis Supervisor ConcurrencyCalibrates supervisor task counts, worker task slots, and heap/off-heap direct buffers aligned with streaming partitions, ensuring zero-lag ingestion and seamless segment handoffs.Overallocates supervisor tasks beyond physical MiddleManager worker slots, leaving real-time tasks queued in pending state and falling behind Kafka retention.Implements custom cron-based micro-batch scripts over HTTP instead of native streaming supervisors, causing frequent data loss and ingestion stalls.Assigns arbitrary maxRowsInMemory thresholds, causing premature intermediate segment spilling to disk and frequent OutOfMemoryErrors on indexing workers.
Cloud Deep Storage Lifecycle (S3/GCS/Azure Blob)Configures reliable cloud deep storage integrations (S3/GCS/Azure Blob) with lifecycle archive tiers, atomic segment push/pull mechanics, and resilient Coordinator kill-task cleanup policies.Fails to automate segment deletion or kill-tasks in deep storage, leaving orphan segment files accumulating hundreds of terabytes in cloud object storage.Mounts shared NFS or local POSIX storage as deep storage, creating single-point-of-failure storage bottlenecks and cluster-wide crash risks.Deploys Druid without durable deep storage configuration, risking catastrophic cluster-wide data loss upon worker node restarts.
Enterprise Security, TLS & Kerberos/RBACImplements end-to-end mTLS encryption across all Druid tiers, Kerberos/OIDC authentication, role-based access control (RBAC), and SQL query-level column masking.Configures basic plaintext HTTP internal communication, leaving cluster control APIs and inter-node segment replication exposed on local networks.Relies on external edge reverse proxies alone without configuring internal Druid authenticator extensions or role-based datasource permissions.Disables authentication extensions entirely to bypass configuration complexity, exposing the Druid Coordinator and Router consoles to unauthorized access.

Apache Druid Engine Failure Modes

Critical Apache Druid Outage Modes We Eliminate

High-throughput streaming Apache Druid clusters encounter severe availability and latency risks when small segments explode Historical server memory, Coordinator locks freeze segment handoffs, or MiddleManager workers exhaust JVM heaps. Our DBREs resolve these breakdown modes:

P1 Critical

Small Segment Explosion Saturating Historical Server Memory

Ingestion pipelines generating millions of uncompacted sub-100MB segments force Historical nodes to maintain massive segment metadata trees in heap. Heap exhaustion triggers prolonged stop-the-world GC pauses, Coordinator assignment lags, and cluster-wide Broker query timeouts.

JusDB Engineering Mitigation:

JusDB DBREs calibrate automated Coordinator compaction rules to consolidate small fragments into 400-600MB segments, optimize partition sharding, and tune Historical off-heap direct memory buffers for optimal scan vectorization.

P1 Critical

Coordinator Metadata Lock Contention Freezing Segment Handoffs

Aggressive metadata polling combined with heavy real-time publishing tasks causes table-level lock contention on PostgreSQL/MySQL metadata stores. Coordinator fails to acknowledge segment handoffs from MiddleManagers, causing streaming tasks to stall and memory buffers to overflow.

JusDB Engineering Mitigation:

JusDB optimizes metadata database connection pooling, adjusts druid.coordinator.period and handoff poll intervals, and decouples segment management loops to ensure non-blocking segment handoffs.

P2 High

MiddleManager Worker Heap Exhaustion Crashing Indexing Tasks

High-cardinality dimension bursts or improperly configured maxRowsInMemory force MiddleManager worker peons to exceed physical JVM heap ceilings during ingestion rollup. Tasks crash with OutOfMemoryError, forcing supervisor re-tries and driving up Kafka consumer lag.

JusDB Engineering Mitigation:

JusDB DBREs tune druid.indexer.fork.property JVM arguments, calibrate maxBytesInMemory thresholds, and configure proactive worker heap monitoring to guarantee smooth indexing handoffs.

Telemetry Runbooks · Non-Blocking Apache Druid Diagnostics

Our Apache Druid DBREs execute non-blocking telemetry inspections to verify segment size distributions, compaction health, and streaming supervisor states without interrupting analytical query workloads:

Druid: Segment Sizing Distribution & Row Count Audit
SQL · sys.segments

Queries the Druid system catalog to evaluate segment count, average segment size in megabytes, and average row count grouped by datasource to detect small-segment fragmentation.

curl -s -X POST "http://localhost:8888/druid/v2/sql" \
  -H "Content-Type: application/json" \
  -d '{"query": "SELECT datasource, COUNT(*) AS num_segments, AVG(\"size\")/1048576 AS avg_size_mb, AVG(num_rows) AS avg_rows FROM sys.segments WHERE is_active = 1 GROUP BY 1"}'
Druid: Kafka Supervisor & Indexing Task Status Forensics
HTTP · Supervisor API

Inspects active Kafka and Kinesis indexing supervisors, reporting status states, task health, generation IDs, and consumer lag offsets across streaming pipelines.

curl -s -X GET "http://localhost:8888/druid/indexer/v1/supervisor"

FAQ

Druid consulting — common questions

Ready to make the call on Druid?

Book a 30-minute scoping call. We'll tell you which engagement shape fits and what the deliverable will look like — before any statement of work.

Related Apache Druid Services

Explore more ways our Apache Druid experts can help with your database infrastructure.