Free audit

View Audit Scope

Failover incidents — sound familiar?

  • ▸ Sentinel quorum problems — 2-node Sentinel with only 1 healthy can't promote anything; you discovered min-replicas-to-writeisn't tuned correctly only when production failed over.
  • ▸ Replica promoted with stale data — a lagging replica fell behind and was elected primary on failover because min-replicas-max-lag wasn't enforced, silently dropping the most recent writes.
  • ▸ Replica failover taking 30s+ — Sentinel takes 10s to detect, 5s to vote, 15s to reconfigure replicas; app sees 30+ seconds of timeouts during what should be a graceful failover.

JusDB HA consultants own the failover playbook + 15-minute incident SLA. Book an HA architecture review →

Single-primary + Sentinel — not multi-shard

Valkey High Availability

In short: Valkey high availability (single-primary + replicas + Sentinel) involves quorum-based Sentinel deployment across AZs, sub-second replica lag monitoring, split-brain prevention via min-replicas-to-write and failover-timeout tuning, and automated failover in 15–30 seconds — plus cross-region async replicas and RDB snapshots for disaster recovery beyond local HA.

Production Sentinel quorum design, replica lag monitoring, split-brain prevention, and 15–30 second automated failover SLAs. For horizontal multi-shard scaling, see Valkey Cluster.

Executive Direct Answer · Valkey High Availability Heuristic

JusDB delivers enterprise Valkey high availability architectures achieving 99.999% uptime through multi-AZ Sentinel quorum configurations and distributed cluster sharding. Our certified DBREs eliminate split-brain write corruption via min-replicas-to-write safeguards, automate sub-5-second primary failover, and conduct quarterly chaos drills, guaranteeing business continuity backed by contractual 15-minute Sev-1 response SLAs.

Uptime: 99.999% Availability SLA·Failover: Sub-5s Automated RTO·Safeguards: Split-Brain Mitigation·Topology: Multi-AZ Sentinel Quorum·Compliance: ISO 27001 & SOC 2

The Coverage

Production HA capabilities

Sentinel Quorum Design

3 or 5 sentinel deployment, quorum/majority sizing, sentinel-monitor configuration, anti-affinity across AZs.

Replication Lag Monitoring

Sub-second lag tracking, master_link_status alerting, replication-offset deltas, lag SLO enforcement.

Split-Brain Prevention

min-replicas-to-write + min-replicas-max-lag tuning, partition-tolerance config, failover-timeout to prevent thrashing.

Automated Failover

down-after-milliseconds tuning, parallel-syncs config, failover SLA tracking, client-side reconnect strategy review.

Health & Status Observability

Prometheus exporter setup, Sentinel & primary dashboards, alert routing for replica-out / split-brain / failover events.

Cross-Region DR

Cross-region async replica placement, snapshot offsite, DR-runbook engineering, RPO/RTO target validation drills.

The Reference Shape

A typical Valkey HA deployment

The shape we deploy by default unless something in the workload pushes us to cluster mode.

Topology

1 primary + 2 replicas (one per AZ) — survives any single-node or AZ outage.
3 Sentinel instances co-located with application servers, spread across the same 3 AZs.
Cross-region async replica for DR (RPO ~30s, manual promotion).

Key parameters

down-after-milliseconds: 10000-15000
min-replicas-to-write: 1 (≥1 connected replica required)
min-replicas-max-lag: 10 seconds
failover-timeout: 180000 (prevents thrashing)
Comparative Matrix · Valkey High Availability Topologies

How JusDB Valkey HA compares to alternative approaches.

Misconfigured Sentinel quorum, unmonitored replication buffers, and lack of split-brain write fencing trigger catastrophic downtime. Here is how our certified DBRE methodology compares across core evaluation vectors:

Evaluation Vector
JusDB DBRE
In-House DBALegacy AgencyDeveloper Generalist
Sentinel Quorum & Multi-AZ Monitor Placement3 or 5-node odd-quorum Sentinel mesh distributed across independent availability zones with strict anti-affinity rules, co-located application monitoring, and zero single-point-of-failure topologies.Even-numbered Sentinel nodes deployed on same host or single AZ; network partition induces 50/50 vote split and total election deadlock.Single Sentinel instance monitoring cluster; any Sentinel crash or transient partition completely disables automated cluster failover.Sentinels deployed on same compute instances as Valkey nodes; node crash simultaneously eliminates monitor and quorum.
Asynchronous Replication Lag & Buffer SizingSub-second lag tracking via master_link_status and replication-offset deltas, paired with calibrated repl-backlog-size and client-output-buffer-limit tuning to prevent full resync loops.Default replication backlog buffers triggering full resync disk thrashing whenever network jitter exceeds 50ms.Ignores replication lag metrics until lagging replicas get promoted with massive data staleness during unexpected primary failures.Unmonitored async replication causing silent write divergence and unbounded memory growth under sustained load.
Split-Brain Prevention (min-replicas-to-write)Enforced min-replicas-to-write and min-replicas-max-lag constraints on primary nodes, isolating partitioned primaries and blocking phantom writes during network splits.Unconfigured min-replicas-to-write; partitioned primary continues accepting isolated writes that are permanently wiped upon healing.Misunderstands partition semantics; relies on client timeouts rather than engine-level fencing to prevent dual-primary states.Completely disables replica write guards, causing catastrophic data corruption when both partitions accept concurrent conflicting keys.
Automated Sub-5s Primary Failover & Client DiscoveryCalibrated down-after-milliseconds, fast Sentinel election, parallel replica resync, and client pub/sub +switch-master reconnection protocols achieving sub-5s failover.Default 30s failure detection combined with unhandled client connection caching resulting in 60s+ application write outages.Requires manual administrative intervention and manual DNS updates to redirect traffic following primary host failure.Hardcoded IP endpoints in client configuration; failover leaves application persistently attempting writes to dead node.
Cross-Region Disaster Recovery & RPO GuaranteesEngineered cross-region asynchronous replication standby, automated point-in-time RDB/AOF archival to object storage, and contractual RPO < 30s.Daily snapshot backups stored locally without cross-region replication, resulting in 24-hour RPO data loss during regional outages.Manual off-site tar exports with unverified restore procedures; RTO exceeds 12 hours during catastrophic site failure.No off-site backup or cross-region DR plan; single cloud region failure results in total, irreversible data loss.
Chaos Engineering & Sentinel Leader Election DrillsQuarterly automated chaos drills simulating network partitions, hard node kills, and split-brain scenarios to validate quorum election and SLA bounds.Untested failover assumptions; first real failover test occurs during an unannounced 3 AM production crisis.Theoretical runbooks on paper without empirical failure testing; election timeouts fail under production connection stress.Zero chaos testing or disaster recovery drills; panic-driven ad-hoc reconfiguration during infrastructure incidents.

High Availability Failure Modes

Critical Valkey HA Failure Modes We Prevent

Even-node Sentinel deadlocks, un-fenced network partitions, and replication lag spikes threaten high availability. We engineer resilience into every tier to eliminate these failure modes:

P1 Critical

Split-Brain Dual-Primary Writes During Network Partition

When an isolated primary becomes disconnected from Sentinel monitors but remains reachable by a subset of application clients, it continues accepting writes. Once Sentinel elects a new primary and the network heals, the old primary is demoted to a replica and its isolated writes are permanently truncated.

JusDB Engineering Mitigation:

JusDB DBREs enforce min-replicas-to-write 1 and min-replicas-max-lag 10 on all primary nodes, automatically converting partitioned primaries to read-only before failover occurs.

P1 Critical

Sentinel Quorum Loss in Even-Numbered Node Topologies

Deploying an even number of Sentinel instances (such as 2 or 4 nodes) or co-locating them in only two availability zones leads to quorum starvation during network partitions. Sentinels cannot reach majority consensus, completely blocking automated failover elections.

JusDB Engineering Mitigation:

JusDB enforces an odd-numbered Sentinel mesh (minimum 3 or 5 nodes) distributed across independent AZs with anti-affinity scheduling, guaranteeing quorum survival during single-AZ outages.

P2 High

Asynchronous Replication Lag Exceeding Client Read Durability

Write-intensive bursts can exhaust the replication backlog buffer (repl-backlog-size), causing replicas to lag hundreds of megabytes behind. If an unmonitored lagging replica is promoted, application reads experience severe data regression and missing session states.

JusDB Engineering Mitigation:

JusDB deploys real-time offset delta alerting, sizes repl-backlog-size to absorb peak multi-hour churn, and configures Sentinel failover selection priorities based on replication offset freshness.

Telemetry Runbooks · Non-Blocking High Availability Diagnostics

Our DBREs monitor Sentinel quorum health, election status, and replica lag metrics using these non-blocking diagnostic commands:

Valkey: Sentinel Quorum & Master State Verification
CLI · Sentinel Quorum

Inspects monitored master health, quorum count, voting validity, and reachable Sentinel nodes via the Sentinel management port.

# 1. Verify Sentinel master state and quorum validity
valkey-cli -p 26379 sentinel masters
valkey-cli -p 26379 sentinel ckquorum mymaster
Valkey: Replication Lag & Slave Offset Telemetry
CLI · Replica Lag

Audits connected slave offsets, link status, and node role to verify continuous sub-second replication synchronization.

# 2. Inspect replication lag, offset deltas, and current node role
valkey-cli info replication
valkey-cli role

FAQ

HA FAQ

Build your Valkey HA topology

Send us your current Valkey/Redis HA setup (or none at all). We'll come back with a sized topology proposal, Sentinel quorum recommendation, and DR runbook outline within 48 hours.

Related Valkey Services

Explore more ways our Valkey experts can help with your database infrastructure.