Why a message queue?

A spiky producer, a slow consumer. Call it directly and the producer dies with it — buffer it in a queue and the spike becomes a backlog instead.

1
Part 1 · Watch the concept

The idea in 21 seconds

A fixed animation — no controls, nothing to configure. It shows the shape of the problem and the shape of the fix, with the story written out underneath.

Loading player…

~21s animated walkthrough: a checkout API blocking on a slow receipt worker vs. the same work buffered in a queue.

The story the animation tells

Checkout works. You press pay, and before the response comes back the API has already rendered a receipt PDF and handed it to the mail server. Two seconds, maybe three on a bad day, and the customer sees a confirmation page with the receipt already on its way. Nobody designed it that way on purpose — rendering the PDF was five lines, and it was simply easier to do it right there in the request.

Then a sale starts. Ten checkouts a second against a worker that renders two. The eleventh request doesn't fail, which is the part that fools you — it waits. So does the twelfth. Each one is holding a request thread hostage while the mail server takes its time, and the pool is only eight threads deep. Within a few seconds every thread in the API is parked inside the receipt code, and the next customer to press pay doesn't get a slow checkout, they get a 503. The cart page is down. Nothing is broken, and everything is unavailable.

The cruelty is that the receipt was never urgent. Nobody refreshes their inbox during checkout. You coupled the thing that must answer in 200ms to the thing that takes two seconds, and the slow one won.

So the video flips it: a queue slides between them. Now the API writes a message and returns — three milliseconds, no matter what the mail server is doing — and the workers pull jobs when they're ready. The spike doesn't vanish. Watch the slots fill: ten deep, and the oldest receipt is now three seconds behind. But a backlog is a number on a dashboard you can add consumers to, where an exhausted thread pool was a checkout page that wouldn't load.

The point. A queue doesn't make the work faster — the same two jobs a second still trickle through. What it buys you is that the producer's latency stops depending on the consumer's throughput. You trade an instant, visible failure for a delay you have to monitor, which is why the queue depth graph is the first dashboard you build after adding one.

2
Part 2 · Run the numbers

Interactive producer/consumer simulator

Now drive it yourself. Pick a mode, set the produce rate and consumer capacity, then run or spike the load and watch real counters — producer latency, queue depth, lag, 503s, timeouts — respond to your settings.

Blocking call
The API waits for the worker on every request, so its latency is the worker's latency. Sustainable while 3/s stays under 4/s.
produced
0
accepted
—
503s
0
timed out
0
Producer-visible latency
—
the caller waits for the whole job, so this tracks the worker
End-to-end completion
—
same number — the customer is holding the connection for it
Request thread poolt = 0.0s
0 / 12 threads held · oldest waiting 0.0s of 2s timeout
Consumer capacity
4/s
2 × 2 jobs/s · 0 completed · demand 0.8× capacity
0% of consumers busy
Last decisions
Run traffic to see decisions.
Why this matters
  • Latency stops being shared. In sync mode the producer's latency is the consumer's latency. Queued, the write costs 3ms no matter how far behind the workers are — compare the two latency tiles above.
  • The failure changes shape, not existence. Spike sync mode and you get 503s and timeouts on the checkout path. Spike queued mode and you get depth and lag — the same overload, moved somewhere you can watch it and add consumers.
  • A queue is not more capacity. Keep the produce rate above consumer capacity and queued mode never recovers: depth climbs until it hits the cap and the producer gets back-pressured anyway. Buffers absorb bursts, not deficits.
  • Lag is the new SLO. "Oldest message behind" is what your customer actually feels once the response no longer carries the work — which is why queue depth is the first dashboard you build after adding one.
Traffic console
Run steady traffic or fire a spike and compare modes.
Try this
Hit “Send spike” and watch the thread pool fill right to the cap — callers start timing out at 2s and new checkouts get a 503, even though the worker never crashed and is still draining at full speed. Then switch to Queued and fire the same spike.