Why replication?
Reads and writes on one node, then reads moved out to replicas. Watch writes stop starving — and stale reads start.
The idea in 22 seconds
A fixed animation — no controls, nothing to configure. It shows the shape of the problem and the shape of the fix, with the story written out underneath.
~22s animated walkthrough: one primary serving every read and write vs. writes on the primary and reads on async replicas.
One database, and for a long time that is the right answer. The product page reads from it, the checkout writes to it, and both are comfortably inside what one box can do. Nothing about the design is wrong yet.
Then you get written about. Traffic is ninety-five percent reads — people browsing, refreshing, sharing links — and the read rate goes up four-fold in an afternoon. The reads get slower, which you expected. What you didn't expect is the page that stops working: checkout. Because reads and writes are queued against the same CPU, and when forty browse queries are waiting their turn, the write that records a payment is just item forty-one in the line. The thing earning money is now starving behind the thing that is free.
So you add read replicas. Writes stay on the primary — which is suddenly bored, because writes were only ever a small fraction of the traffic — and reads fan out across copies that you can keep adding. Checkout gets its 3ms back, browse latency flattens, and the graph looks fixed.
Except replication is asynchronous. The video shows the primary committing v41 while a replica is still answering from v39, and that is not a bug you can configure away: it's the cost of the copy existing somewhere else. A user updates their profile and the next page load shows the old name. Nothing is broken, nothing errored, and the data is simply from a moment ago.
The point. Replication buys read capacity with consistency, and the exchange rate is set by your write rate times your replication delay — not by how many replicas you add. The fix for the stale read isn't a bigger fleet, it's deciding which reads are allowed to be slightly old, and routing the ones that aren't back to the primary.
Interactive read/write simulator
Now drive it yourself. Pick a topology, set the read and write rates, then spike the reads and watch real counters — read and write latency, rejections, replication lag, stale reads — respond to your settings.
- A read spike is a write outage. Spike single-primary mode: the write rate never changed, but write latency tracks the read queue because they share one node. That is the failure replication actually fixes.
- Replicas scale reads, not writes. Add replicas and read capacity grows; the write path is still exactly one node. If writes are your bottleneck, this pattern does nothing for you — that is what sharding is for.
- Staleness is priced by writes × delay. The lagging-read count jumps when you raise the write rate or the replication delay, and does not move at all when you add replicas. Note what the counter means: the replica was behind the primary when it answered — not every one of those reads touched the changed row, but any of them could have. With writes arriving continuously, an async replica is essentially never fully caught up, which is what "eventual" actually looks like.
- The user who wrote is the one who notices. They submitted the write that is still in flight, so their next read is the most likely to be stale — which is why read-your-writes routes that one back to the primary.
Related guides & quizzes
The simulator builds the intuition. These go deeper on the tradeoffs, then check whether it stuck.