Production DBA Comparison
DynamoDB vs Cosmos DB
Choose Amazon DynamoDB if your infrastructure is native to AWS, requiring single-digit millisecond key-value lookups, serverless auto-scaling, and predictable RCU/WCU billing with DAX microsecond caching. Choose Azure Cosmos DB if hosted on Azure, demanding five tunable consistency levels, multi-API wire compatibility, or multi-master active-active writes across global regions.
RCU/WCU vs RU/s. Global Tables vs multi-master. Single-API vs five-API surface. Vector search, consistency models, partition-key design — the cloud-native NoSQL decision in 2026.
Sound familiar?
- ▸ RCU/WCU provisioning math doesn't match the real load shape — on-demand is doubling spend but provisioned mode is throwing 400s during peak, and the partition-key design is suspect.
- ▸ RU/s estimation for a new Cosmos workload is wrong by 5x and the autoscale tier hasn't saved enough — finance wants a rebid before the next quarter, and the team needs a real capacity model.
- ▸ Multi-region requirements just landed — regulator wants in-region writes, app team wants active-active. Global Tables vs Cosmos multi-master is the architecture call that hasn't been made.
JusDB consultants build the written DynamoDB-vs-Cosmos DB decision with a partition-key audit attached. Book a cloud-NoSQL scoping call →
Architectural Analysis
DynamoDB vs Cosmos DB — Comparative Evaluation Matrix
Evaluate the six core technical vectors separating AWS DynamoDB key-value partitioning from Azure Cosmos DB multi-model Request Unit architectures, mapped alongside JusDB production reliability engineering.
| Evaluation Vector | Amazon DynamoDB | Azure Cosmos DB | JusDB DBRE Architecture |
|---|---|---|---|
| Architecture & Storage Subsystem | SSD-backed distributed storage partitioned into 10GB physical partitions. Automatic partition splitting triggered at >10GB storage or throughput limits (1,000 WCU / 3,000 RCU). | Log-structured ARS (Atom, Record, Sequence) engine on physical partitions up to 50GB. Multi-model abstraction layer exposing NoSQL, MongoDB, Cassandra, Gremlin, and Table APIs. | Partition key entropy audits, synthetic key salting algorithms, multi-tenant physical partition distribution modeling, and hot-partition isolation frameworks. |
| Concurrency, Throughput & Latency Profile | Single-digit millisecond latency (1–3ms, microsecond with DAX caching). Rigid per-partition throughput caps (1,000 WCU / 3,000 RCU) trigger HTTP 400 throttle exceptions on hot keys. | Sub-10ms p99 SLA for 1KB reads/writes globally. Throughput provisioned in Request Units (RU/s); partition saturation triggers HTTP 429 Too Many Requests responses. | Distributed client-side token bucket rate limiting, jittered exponential backoff tuning, DAX/Redis read-through caching tiers, and autoscale RU/s threshold calibration. |
| Failover, High Availability & RTO | Regional multi-AZ synchronous replication across 3 AZs (99.99% SLA). Global Tables provide multi-region active-active replication with asynchronous last-writer-wins (LWW) conflict resolution. | 99.999% SLA across multiple regions with multi-master writes. 5 tunable consistency levels (Strong, Bounded Staleness, Session, Consistent Prefix, Eventual) trading latency for guarantees. | Multi-region failover runbooks, custom conflict resolution registration, replication lag monitoring across cloud regions, and cross-cloud disaster recovery pipelines. |
| Cost Structure & Billing / Resource Utilization | Billed per RCU/WCU (Provisioned) or per request (On-Demand: $1.25/M writes, $0.25/M reads) plus storage ($0.25/GB/mo). Spiky unindexed workloads cause steep On-Demand cost surges. | Billed per provisioned RU/s ($0.008/hr per 100 RU/s), autoscale RU/s (10x dynamic range), or Serverless ($0.25/M RUs) plus storage ($0.25/GB/mo). High indexing write RU overhead. | FinOps capacity modeling: RCU/WCU vs RU/s utilization profiling, automated provisioned vs on-demand break-even calculators, and index pruning reducing unnecessary write charges. |
| Operational Overhead & DBA Maintenance | Zero OS or engine patching; serverless maintenance. Operational effort centers on partition key design, GSI cardinality management, TTL pruning, and IAM access controls. | Zero server management. Operational complexity lies in indexing policy tuning (selective inclusion/exclusion), partition key path selection, and RU consumption optimization. | 24/7/365 DBRE telemetry: continuous partition skew detection, automated indexing policy optimization, TTL expiration verification, and sub-15m P1 incident response. |
| Ecosystem, Tooling & Migration Path | Tight AWS ecosystem integration (DynamoDB Streams, Kinesis, Lambda, EventBridge, IAM). Migration off DynamoDB requires query translation and application-tier rewrites. | Azure-native integration (Azure Functions, Event Hubs, Synapse Link). Multi-API wire compatibility eases initial migration from MongoDB/Cassandra, though behavioral nuances exist. | Automated cross-cloud CDC pipelines (DynamoDB Streams to Kafka/Debezium), schema translation harnesses, dual-write shadow validation, and zero-downtime cutover orchestration. |
Resilience Engineering
DynamoDB & Cosmos DB Production Failure Modes
Critical cloud-native NoSQL failure scenarios analyzed and remediated by JusDB DBREs to prevent partition throttling, indexing RU cost runaway, and GSI backpressure write freezes.
Hot Physical Partition Throttling Under Skewed Hash Keys
A poorly distributed partition key concentrates queries onto a single 10GB physical partition in DynamoDB (exceeding 1,000 WCU or 3,000 RCU) or a Cosmos DB partition (exceeding provisioned RU/s). CloudWatch or Azure Monitor logs continuous HTTP 400 / 429 throttle exceptions while global cluster capacity sits idle at <15% utilization.
Implement synthetic key salting (suffixing keys with randomized hashes), decouple write surges with SQS/Event Hubs queuing, and deploy JusDB hot-key monitoring to dynamically rebalance shard allocations.
Cosmos DB Index Over-Provisioning and RU Burn Escalation
Cosmos DB default indexing policies automatically index every document path, string, and numeric property. On high-velocity ingestion workloads (e.g., IoT telemetry or financial logs), write operations burn 15–25 RU/s per insert instead of the expected 5–6 RU/s, exhausting provisioned autoscale budgets and doubling monthly Azure invoices.
Audit document schema paths using Azure Cosmos DB metrics, transition to explicit index inclusion/exclusion policies, strip spatial/composite indexes from write-heavy collections, and isolate archival telemetry.
DynamoDB Global Secondary Index (GSI) Backpressure Throttling
When a GSI has lower provisioned write capacity than the base DynamoDB table, write operations to the base table are throttled (ProvisionedThroughputExceededException) even if the base table has ample headroom. DynamoDB enforces backpressure from unscaled GSIs to maintain index synchronization.
Configure synchronized auto-scaling policies between base tables and all attached GSIs, implement adaptive capacity thresholds, and isolate low-cardinality query patterns into sparse GSIs.
Telemetry & Observability
Production Diagnostic Runbooks
Non-blocking inspection commands executed via AWS CLI and Azure CLI to audit partition key skew, consumed capacity units, Request Unit charges, and throttle exception rates.
Audits consumed read/write capacity units, detects partition throttle events, and identifies GSI backpressure limits.
# 1. Query CloudWatch for throttled requests on base table and GSIs
aws cloudwatch get-metric-data \
--metric-data-queries '[
{
"Id": "m1",
"MetricStat": {
"Metric": {
"Namespace": "AWS/DynamoDB",
"MetricName": "ThrottledRequests",
"Dimensions": [{"Name": "TableName", "Value": "prod_orders"}]
},
"Period": 60,
"Stat": "Sum"
}
}
]' \
--start-time "$(date -u -v-1H +%Y-%m-%dT%H:%M:%SZ)" \
--end-time "$(date -u +%Y-%m-%dT%H:%M:%SZ)"
# 2. Check table description for GSI capacity status and partition metrics
aws dynamodb describe-table \
--table-name prod_orders \
--query "Table.{Status:TableStatus,ItemCount:ItemCount,SizeBytes:TableSizeBytes,BillingMode:BillingModeSummary.BillingMode,GSIs:GlobalSecondaryIndexes[*].{IndexName:IndexName,Status:IndexStatus,RCU:ProvisionedThroughput.ReadCapacityUnits,WCU:ProvisionedThroughput.WriteCapacityUnits}}"Measures normalized RU consumption per physical partition, detects 429 response rate spikes, and checks autoscale ceilings.
# 1. Audit normalized RU consumption percentage across partition key ranges
az monitor metrics list \
--resource "/subscriptions/{subId}/resourceGroups/{rg}/providers/Microsoft.DocumentDB/databaseAccounts/{account}" \
--metric "NormalizedRUConsumption" \
--interval PT1M \
--aggregation Maximum
# 2. Check 429 throttle rate and autoscale throughput allocation
az cosmosdb sql container throughput show \
--account-name prod-cosmos-cluster \
--resource-group prod-data-rg \
--database-name core_db \
--name orders \
--query "{CurrentRU:resource.throughput,AutoscaleMaxRU:resource.autoscaleSettings.maxThroughput}"When DynamoDB wins
- You're on AWS and the rest of the stack is AWS-native.
- Steady, predictable read/write throughput — provisioned capacity is cheaper.
- Access pattern is genuinely key-value — no need for multi-API flexibility.
- You want the deepest AWS integration: Lambda triggers, Streams, IAM, KMS, VPC.
- PITR + Global Tables cover your DR and multi-region needs.
- You're happy with eventual consistency cross-region (last-writer-wins).
When Cosmos DB wins
- You're on Azure and the rest of the stack is Azure-native.
- Tunable consistency per workload is a real architectural requirement.
- You want MongoDB / Cassandra / Gremlin API to migrate without rewriting.
- Vector search co-located with primary documents simplifies RAG architecture.
- Multi-master with strong cross-region consistency is required.
- Automatic indexing on all properties is operationally easier than DynamoDB's GSI/LSI design.
Migration
Migration paths between DynamoDB and Cosmos DB
DynamoDB → Cosmos DB
Pick the Cosmos API first — usually the native SQL API for a fresh design or Table API for the closest DynamoDB ergonomic match. Re-shape the data model, write the migration script (DynamoDB scan → Cosmos batch insert), and validate the new partition key.
Cosmos DB → DynamoDB
Trickier when you're moving off Cosmos MongoDB API — re-targeting at DynamoDB native requires rewriting all queries since DynamoDB has no $match / $group / $lookup. Often the easier path is Cosmos → MongoDB Atlas instead.
Either → MongoDB Atlas (multi-cloud)
Common destination when teams want out of single-cloud lock-in. Atlas runs on AWS, Azure, and GCP equally well, with replica sets and sharding native. We help model the migration plus the operational handoff.
Common questions
Need a written DynamoDB-vs-Cosmos decision?
We audit the partition-key design, model the throughput, and write the migration runbook — for either direction or for the "move off cloud-native" option.