Production DBA Comparison
Aurora MySQL vs RDS for MySQL
Choose Aurora MySQL for read-heavy workloads requiring up to fifteen low-lag auto-scaling replicas, sub-sixty-second failover, or Serverless v2 elasticity. Choose RDS MySQL for cost-sensitive, write-heavy, or steady OLTP workloads needing full MySQL 8.4 engine parity without paying Aurora's 20–50% compute premium or unpredictable per-I/O billing charges.
Aurora and RDS for MySQL are both AWS-managed services, but they're architecturally very different beasts. Aurora uses a shared-storage cluster model with 6-way replication and decoupled compute. RDS for MySQL is a closer-to-vanilla MySQL with EBS storage. The right pick depends on your read-write ratio, scale, HA requirements, and how much per-IO cost you can stomach. This guide is the production-DBA view of when each wins.
Aurora vs RDS — sound familiar?
- Aurora bill higher than expected — I/O-optimised vs standard, instance class, multi-AZ — the cost model has 3+ axes you didn't model. The bill doubled in 6 months.
- 5x performance claim doesn't materialise — Marketing said 5x. Your workload sees 1.2x. Workload is write-heavy and Aurora's read-replica architecture doesn't help you.
- Stuck on RDS, wondering about Aurora — Considering the migration but unclear whether the throughput jump justifies the cost. Need the actual decision math, not vendor brochure.
JusDB DBAs run both in production. We'll give you the honest answer in 30 minutes — no vendor pitch. Book a comparison call →
Architectural Analysis
Aurora MySQL vs RDS for MySQL — Comparative Evaluation Matrix
Examine the six critical technical vectors separating AWS Aurora's distributed log storage from traditional Amazon RDS EBS architectures, mapped alongside JusDB production reliability engineering.
| Evaluation Vector | Aurora MySQL/Postgres | Amazon RDS | JusDB DBRE Architecture |
|---|---|---|---|
| Architecture & Storage Subsystem | Distributed log-structured shared storage across 3 AZs (6 copies). Compute decoupled from storage with quorum writes (4/6) and reads (3/6). | Monolithic compute coupled to EBS volumes (gp3 or io2 Block Express). Storage tied to specific instance attachment and EBS volume limits. | Hybrid architectural engineering matching storage engines to IOPS intensity, automated volume expansion, NVMe caching tiers, and multi-cloud portability. |
| Concurrency, Throughput & Latency Profile | Up to 3-5x read throughput via up to 15 auto-scaling replicas sharing the cluster volume with near-zero lag (<20ms). Single-writer write concurrency bottlenecked by cluster primary. | Read throughput scaled via async binlog/streaming replicas with variable lag (seconds to minutes) under heavy writes. Up to 5-15 read replicas bounded by EBS IOPS limits. | Layer-7 query routing (ProxySQL/PgBouncer) with connection multiplexing, statement-level read/write segregation, and query cache layer eliminating hot-row lock queues. |
| Failover, High Availability & RTO | Rapid automated failover (typically 30–60s) via cluster reader promotion without storage re-attachment. Aurora Global Database offers cross-region lag <1s. | Multi-AZ synchronous block-level standby failover (60–120s) requiring DNS endpoint flip, crash recovery replay, and EBS remounting. | Zero-downtime consensus failover with client-side connection draining, automated VIP re-pointing, and orchestrated split-brain fencing achieving sub-15s RTO. |
| Cost Structure & Billing Predictability | Complex multi-tier billing: compute instance + storage ($0.10/GB) + per-request I/O charges ($0.20/1M IOs) or 30% storage premium on I/O-Optimized. High risk of surprise surges. | Predictable flat pricing based on compute hours and provisioned gp3 storage/IOPS. No separate per-I/O billing charges regardless of query frequency. | Comprehensive FinOps workload modeling: I/O audit, right-sizing instance classes, I/O-Optimized vs Standard break-even calculation, reducing cloud database spend by 30–50%. |
| Operational Overhead & DBA Maintenance | Managed OS and patching, automated point-in-time recovery and continuous storage auto-scaling up to 128TB. Requires deep knowledge of ACU thresholds and buffer cache mechanics. | Managed infrastructure, automated snapshots, and minor engine upgrades. Storage auto-scaling requires manual step configuration; vacuum/purge overhead remains. | 24/7/365 dedicated DBRE operational support with sub-15m P1 SLA, predictive storage alerts, vacuum tuning, table bloat removal, and proactive parameter group management. |
| Ecosystem, Tooling & Portability | Heavy AWS vendor lock-in. Proprietary storage subsystem cannot be exported or run outside AWS; features often diverge from upstream community MySQL/PostgreSQL versions. | Standard open-source engine compatibility; exportable via native mysqldump, pg_dump, or physical snapshots. Lower vendor switching barriers. | Infrastructure-as-Code database deployment, automated CDC (Debezium/Flink) for zero-downtime cloud exits, and open-source tooling neutrality across AWS, GCP, Azure, and bare metal. |
Resilience Engineering
Aurora & RDS Production Failure Modes
Critical cloud database failure scenarios analyzed and remediated by JusDB DBREs to prevent runaway I/O billing surges, asymmetric failover deadlocks, and EBS storage replication stalls.
Unbounded Aurora Storage I/O Cost Runaway on Unindexed Scans
Full table scans and missing composite indexes on high-traffic tables generate millions of storage subsystem IO requests per minute. On Aurora Standard, this triggers exponential CloudWatch bill spikes where I/O charges exceed the base database compute and storage cost combined.
Conduct continuous Performance Insights statement audits, deploy in-proxy query throttling for unindexed scans via ProxySQL, and transition write/read-heavy clusters to Aurora I/O-Optimized pricing to eliminate per-request I/O metering.
Aurora Reader Promotion Deadlock During Asymmetric Failover
When the primary writer encounters node exhaustion, Aurora promotes a reader based on failover priority tiers. If application connection pools continue hammering the former writer during role negotiation, TCP connection timeouts cascade across services while readers struggle with uncommitted rollback segments.
Enforce RDS Proxy with automated target health detection, configure explicit cluster failover priority tiers (tier-0 for designated failover standbys), and implement aggressive TCP keepalive timeouts on application client drivers.
RDS Multi-AZ Replication Stall Under Write-Intensive EBS IOPS Exhaustion
During massive bulk data loads or concurrent transaction bursts, RDS instances on standard gp3 storage exhaust baseline IOPS or burst credits. Synchronous Multi-AZ block-level EBS replication lags, forcing primary transaction threads into fsync wait states and triggering application database lockups.
Pre-warm and provision dedicated gp3/io2 IOPS and storage throughput allocations ahead of peak loads, leverage asynchronous replication for non-critical reads, or re-architect high-ingestion tables to Aurora's quorum-based log storage layer.
Telemetry & Observability
Production Diagnostic Runbooks
Non-blocking inspection commands executed via CloudWatch metrics, Performance Insights, and MySQL information_schema to audit storage I/O, cache hit ratios, and failover latency.
Measures Aurora storage I/O request rates, buffer pool cache hit ratios, and replica lag across cluster reader instances.
# 1. Audit Aurora buffer pool hit efficiency and read requests mysql -h cluster-endpoint.rds.amazonaws.com -u dbre_admin -p -e " SELECT VARIABLE_NAME, VARIABLE_VALUE FROM performance_schema.global_status WHERE VARIABLE_NAME IN ( 'Innodb_buffer_pool_read_requests', 'Innodb_buffer_pool_reads', 'Innodb_buffer_pool_pages_data', 'Innodb_buffer_pool_pages_total' );" # 2. Inspect live Aurora cluster replication lag across reader hosts mysql -h reader-endpoint.rds.amazonaws.com -u dbre_admin -p -e " SELECT SERVER_ID, LAST_UPDATE_TIMESTAMP, REPLICA_LAG_IN_MSEC FROM information_schema.replica_host_status;"
Identifies thread wait events, synchronous disk fsync latencies, and EBS queue depth saturation on Amazon RDS.
# 1. Audit top wait events and fsync disk latency on primary instance mysql -h rds-master.rds.amazonaws.com -u dbre_admin -p -e " SELECT EVENT_NAME, COUNT_STAR, ROUND(SUM_TIMER_WAIT/1000000000, 2) AS wait_ms FROM performance_schema.events_waits_summary_global_by_event_name WHERE COUNT_STAR > 0 ORDER BY SUM_TIMER_WAIT DESC LIMIT 8;" # 2. Check active long-running transactions and lock wait chains mysql -h rds-master.rds.amazonaws.com -u dbre_admin -p -e " SELECT r.trx_id waiting_trx_id, r.trx_mysql_thread_id waiting_thread, b.trx_id blocking_trx_id, b.trx_mysql_thread_id blocking_thread, TIMESTAMPDIFF(SECOND, r.trx_wait_started, CURRENT_TIMESTAMP) wait_seconds FROM information_schema.innodb_lock_waits w JOIN information_schema.innodb_trx b ON b.trx_id = w.blocking_trx_id JOIN information_schema.innodb_trx r ON r.trx_id = w.requesting_trx_id;"
When Aurora MySQL wins
Read-heavy workloads
15 read replicas with shared storage — near-zero lag. Aurora's read scaling is the headline feature.
HA criticality
30-60s failover vs RDS 60-120s. 6-way replication means data loss tolerance is minimal.
Serverless / variable workloads
Aurora Serverless v2 auto-scales ACUs based on load — RDS doesn't have an equivalent.
Cross-region read replicas
Aurora Global Database has sub-second cross-region lag. RDS cross-region replicas are async.
Storage management
Auto-scaling 10GB → 128TB. No DBA intervention to grow storage.
When RDS for MySQL wins
Cost-sensitive deployments
RDS is 30-50% cheaper at small/medium scale. No per-I/O charges. The Aurora premium is real.
Write-heavy workloads
Aurora's shared storage doesn't help if you have a single write hotspot. RDS gets you closer to bare-metal MySQL write throughput per dollar.
Standard MySQL feature parity
Aurora is mostly compatible but has some divergences (MySQL 8.4 not yet supported, certain replication features). RDS is closer to vanilla MySQL.
Multi-AZ simplicity
RDS Multi-AZ is well-understood and operationally simple. Aurora cluster topology has more moving parts.
Lower lock-in
RDS data exports trivially; Aurora's storage layer is AWS-proprietary.
Migration
Migration paths between Aurora MySQL and RDS for MySQL
RDS MySQL → Aurora MySQL
Use the "Create Aurora replica" option in RDS console for near-zero-downtime; then promote. Validate replication lag and connection pool.
Aurora MySQL → RDS MySQL
mysqldump + binlog replay; or DMS task with full + CDC. Test on staging first — some Aurora features won't exist.
Self-hosted MySQL → Aurora/RDS
DMS migration task with CDC for zero-downtime. Aurora preferred for HA-critical; RDS for cost-sensitive.
FAQ
Common questions
Need help deciding?
We run both in production. 30-minute call, honest answer for your specific workload, no vendor pitch.