The shape of the whole thing
Five stages, from "a request failing isn't a fact about the request" to the questions a principal asks before an incident asks them. Read it top to bottom, or jump straight to the part you came for.
Introduction
A retry is the most reasonable thing you can do when a request fails, and the most dangerous, and it's both for exactly the same reason: it doubles down on a request that just failed. Get the timing right and you paper over the small constant rudeness of the network — the dropped packet, the GC pause, the half-second a node spends rebooting. Get it wrong and you take a service that was having a bad minute and hold it under until it has a bad hour.
A request failing is not a fact about the request. It's a fact about this attempt, on this connection, at this instant. The server may have done the work and died before it could answer. It may never have heard you. From where you're standing, every one of these looks identical — silence, then a timeout — and the only honest reply to silence is to ask again. Carefully.
"Carefully" is carrying the entire chapter on its back. Without retries, every transient blip becomes a user-visible error, and in a system where one request fans out to a dozen downstream calls, a hiccup in any one of them fails the whole thing. A payments company that treated every timeout as permanent would decline a card every time a GC pause hit a downstream service. Nobody runs a business that way. Retries are not optional.
And yet: retries amplify load, and they do it at precisely the worst moment. A service fails because it's overloaded; you respond by sending it more of the same traffic that overloaded it, multiplied by every client that just timed out. It's the infrastructure version of telling a drowning person to swim harder. The service struggling under 1,000 requests is now under 2,000 — and you haven't helped it recover. You've finished it off.
Fenced by a circuit breaker (Chapter 4) and made safe to repeat by idempotency (Chapter 2) — none of these come free. The road from "retry on timeout" to "retry without starting a stampede" runs directly through at least one production incident. Most teams pay that toll in person. This chapter is an attempt to let you skip the line.
The Problem, Actually
"Retries cause load" is true and useless. Here's the specific shape of the disaster. Service A calls Service B. B slows down — not crashes, slows, which is worse — because a slow query crept onto the hot path, or a replica fell behind. A's calls start timing out. A retries. Immediately. Because that's what the code does, and the person who wrote the code was not thinking about today.
B was handling 1,000 requests a second. Now it's those 1,000 plus 1,000 retries from the clients that just timed out — 2,000 aimed at a service already buckling at 1,000. B slows further. More clients time out. More retries. Inside a minute you've gone from "B is a little slow" to "B is gone, and the retry storm is the outage now."
The detail that makes this miserable to debug is that cause and effect come apart in the dashboards. The slow query that kicked it off may have cleared in thirty seconds. The retry storm it lit runs for minutes, because the retries are now generating the load that generates the timeouts that generate the retries. The original problem is long dead. The incident is running on its own exhaust — a perpetual-motion machine whose only output is pages. I once spent the first forty minutes of an outage hunting a downstream that had been healthy for thirty-nine of them.
The first fix anyone reaches for is to wait between retries, and it doesn't work for a reason nobody sees coming: all 1,000 clients that timed out run the same code, so they all wait the same second, then all retry at the same instant. You didn't spread the load — you gathered it up, held it for a second, and delivered it as a single synchronized punch. This is the thundering herd, and fixed delays summon it with eerie reliability.
The Naive Solutions & What They Cost
Every engineer who has implemented retries has written approximately this sequence, in this order, usually one production incident apart.
Every step in that sequence is locally reasonable. Each client, considered alone, behaves impeccably. The catastrophe is entirely emergent — it lives in the correlation between a thousand well-mannered clients, none of which can see the other 999. You can read every retry block in the codebase, find nothing wrong with any of them, and still have a storm waiting for the next slow Tuesday. The bug isn't in the code. It's in the arithmetic of everyone running the same correct code at the same moment.
Failure Modes Worth Naming
A hundred clients, each willing to retry ten times, turn 100 requests into 1,000 — ten times the load, aimed at a service whose entire problem was that it already had too much. A healthy-looking 20% of headroom does not survive contact with a 10× multiplier. The math is indifferent to the fact that every individual client was being sensible; sensible, multiplied by a hundred and correlated in time, is a DDoS you built yourself and now pay to host.
Full exponential backoff over five attempts is 1 + 2 + 4 + 8 + 16 = 31 seconds of waiting — a peculiar amount of time to spend when the user's browser gave up at 30. You've built retry logic that structurally cannot succeed before the only person who cared has left the building. The retries fire anyway, because the service has no idea the caller is gone — burning threads and downstream capacity to compute an answer with nowhere to go. It's the office worker still polishing slides for a meeting that ended an hour ago.
A POST /payments returns a 504; the original charge actually succeeded but the response evaporated on the way back; the retry charges the card again. Double charge, over-decremented inventory, a refund nobody's proud of. "Safe to repeat" is not a property you get for free — it's idempotency keys, dedup logic, and a handler that treats at-least-once delivery as exactly-once on purpose. Bolt retries onto a system that never thought about idempotency and you haven't added resilience; you've added a data-corruption bug on a timer.
Amplification with a blast radius. A retries B; while B strains to recover, its own calls to C and D are timing out, so B retries those; meanwhile C and D have other clients also retrying. Every hop multiplies the load on the hop beneath it. The thing that takes down the company was not the service with the original hiccup — it was every service standing behind it, each faithfully amplifying the panic of the one in front.
Pattern 1 · Exponential Backoff with Jitter
The mechanism is the one AWS wrote up in 2015, in a post that remains the clearest thing anyone has published on the subject: wrap the call in a loop, and on each failure sleep for a random interval drawn from zero up to an exponentially growing ceiling, capped so it can't run away. (Appendix A.5)
The word doing the real work is full. "Full jitter" means the wait is uniformly random across the whole window — random(0, window) — not the exponential value with a little noise sprinkled on top. People conflate the two and assume the difference is cosmetic. It isn't: full jitter spreads retries far more evenly at the tail, which is exactly where a recovering service is most fragile. The gap is a few characters of code and a meaningfully gentler load curve.
A thousand clients each making an independent random draw from the same interval will, with no coordination whatsoever, spread themselves smoothly across it. No central scheduler, no service-mesh magic, no clients talking to each other. Each one acts in pure isolated self-interest, and the aggregate is cooperative. It's one of the few places in distributed systems where selfishness composes into something polite, and it costs exactly one call to a random number generator.
Pattern 2 · Retry Budgets
The trouble with max_attempts=5 is that it's a number with no referent. Nobody computed it; somebody picked it. Retry budgets throw out the per-client count and ask a system-level question instead: what fraction of my total traffic am I willing to let be retries? Google's SRE book documents the pattern, and the recommended ceiling is 10% — with the pointed footnote that even 10% is high for a healthy system. (Appendix A.6)
A service taking 100 requests a second with a 10% budget sees on the order of 110 — not 400, not 1,000 — no matter how many individual clients are panicking. And the real ceiling is tighter than that clean number suggests: as the dependency sickens and its successes dry up, the retry ratio crosses the line faster, so the budget clamps down harder at exactly the moment you need it to. The budget doesn't care how many clients there are. It caps the sum.
The math is the smaller half of the value. The larger half is that the budget forces a conversation a team would otherwise never have: what fraction of our requests are we willing to spend on retries? A team that has answered that has started treating retry behavior as a property of the system. A team that hasn't will be introduced to the question in production, by the system, at a time of the system's choosing — and the system has a cruel sense of timing.
Pattern 3 · Deadline-Aware Retries
The enabling mechanism is context propagation: the original caller's deadline rides along through the entire call chain, so every service knows not just that it should hurry but exactly how much time is left. gRPC does this natively; HTTP services do it by convention — a header (X-Request-Deadline) that every hop reads, honors, and passes forward, and that exactly one team always forgets to forward. With a deadline in hand, the retry decision grows a precondition: before sleeping, check whether there's enough time left to bother. (Appendix A.7)
If you've got two seconds left and the next backoff is four, you will certainly miss the deadline — so fail now, immediately, rather than sleeping four seconds to produce an answer the caller abandoned two seconds ago. That check is a small piece of code and a genuine shift in posture: don't retry for your own benefit when there's no longer anyone on the other end to benefit. It kills a whole category of waste where services deep in the graph retry heroically for a request the API gateway gave up on ages ago. It's a kitchen still cooking an order for a table that paid and left.
Trade-offs Worth Arguing About
Retry count vs. amplification
Take a downstream at 70% utilization — 30% headroom, which on paper reads as healthy. A client that retries once on failure sends, at the bad moment, up to 2× its normal volume. 2× into 70% is 140%, which is over the line. Read that twice, because it's the whole trap: the retries are the cause, not the response. The service didn't fail and then get retried. It got retried and then failed.
Three retries for most off-critical-path services; one retry for anything on the critical path. The reasoning is latency, not load — on a critical path, the time spent grinding through backoff cycles can blow past the user's patience, so you're better off failing fast with a clean error than spending ten seconds retrying your way to the same error with worse manners.
Which backoff strategy
Sounds like it should be a long debate; usually isn't. Exponential with full jitter is the right default for anything under load. Linear backoff earns its place only in the narrow case where you care about fairness over spreading — rate-limited clients with a known reset window. Pure exponential without jitter is the one to be suspicious of: it looks like the sophisticated choice and quietly keeps the clients correlated, so it helps less than its reputation promises.
Which errors to retry — the status code is a hint, not a contract
Retrying everything uniformly is how you turn one bad request into the same bad request, sent five times. The policy has to be fitted to the actual error space of the actual service.
What the Companies Actually Built
The Principal Engineer's View
The question that matters most is not "which backoff algorithm should we use?" It's "does anyone here actually know what our retry behavior is doing to our dependencies right now?" At most companies the honest answer is no — retry behavior is buried three layers deep in client libraries, set to a default sensible for the service it shipped with and arbitrary for yours, and surfaced on precisely zero dashboards. The first time anyone looks at it directly is mid-cascade, reverse-engineering it from the wreckage.
Three metrics that deserve the observability you'd never skip for latency
High retry rate, low retry success rate. That isn't resilience working. That's a storm in progress, and the retries aren't helping — they may be the only reason it isn't recovering.
Questions to put to your team, out loud, before an incident does
The last one is the question most teams can't answer with a straight face, because retry behavior is almost always validated in staging, where the load is a rounding error and the thundering herd can't form for lack of a herd. You need enough concurrent clients to stampede before you can watch a stampede, and you only have those in production. Which is the entire case for chaos engineering — because the alternative is letting the system pick the moment, and the system always picks 2 a.m.
Exercise · Keep Within the Budget
Service A sends 100 RPS to Service B. B can sustainably handle 110 — call it 10% headroom. B has a brief degradation. A's current policy: up to 3 retries on any 5xx, fixed 1-second delay between attempts.
The point isn't the answer — it's the argument you have with yourself getting there. Paste it into an AI and you'll get a clean, confident response that skips the only thing that matters: your dependencies, your traffic shape, your blast radius.
Connections to Other Chapters
A retry is never a private decision. It feels like one — a few lines at a single call site, handling a single failure — and that's exactly the illusion that keeps retry storms perennial. The blast radius of those few lines is every service downstream of you, multiplied by every other client running their own reasonable version of the same few lines, correlated by the shared instant you all failed together. Backoff, jitter, budgets, and deadlines aren't four tricks; they're four ways of forcing a local decision to account for a global consequence. Get them right and a thousand clients failing at once produce a survivable drizzle. Get them wrong and they produce a service that was fine until its friends tried to help.
Appendix A · Reference Implementations
The chapter keeps code out of the narrative on purpose. Here it is — collected, runnable, annotated. All of it is illustrative rather than production-hardened: the snippets skip the timeouts, error taxonomy, thread safety, and metrics you'll need in the real thing, and the first four are deliberately broken in instructive ways.
Next: Chapter 6 — Request Coalescing: turning a thousand identical in-flight requests into one piece of work, shared.