Marginalia Systems Under Load / Ch.05 Retries
◇Overview
Systems Under Load — Chapter 05

Retries & Backoff

The most reasonable thing to do when a request fails, and the most dangerous — for exactly the same reason.

27 min read 7 code listings · 5 figures See the concept map ↓
Chapter 05 · Concept map

The shape of the whole thing

Five stages, from "a request failing isn't a fact about the request" to the questions a principal asks before an incident asks them. Read it top to bottom, or jump straight to the part you came for.

Foundations Patterns Trade-offs Case studies Operating it
01
Foundations
Retries are not optional — and the storm is manufactured entirely by the retries.
02
The Core Patterns
Four ways to force a local decision to account for a global consequence.
03
The Trade-offs
Retry count against amplification — counterintuitive until you've run the numbers with your own outage attached.
04
What Companies Built
Google, Stripe, and Netflix arrived at the same discipline through three different doors.
05
Operating It
Does anyone here actually know what our retry behavior is doing right now?
Part III · Failure

Introduction

A retry is the most reasonable thing you can do when a request fails, and the most dangerous, and it's both for exactly the same reason: it doubles down on a request that just failed. Get the timing right and you paper over the small constant rudeness of the network — the dropped packet, the GC pause, the half-second a node spends rebooting. Get it wrong and you take a service that was having a bad minute and hold it under until it has a bad hour.

The sentence that took a few outages to believe

A request failing is not a fact about the request. It's a fact about this attempt, on this connection, at this instant. The server may have done the work and died before it could answer. It may never have heard you. From where you're standing, every one of these looks identical — silence, then a timeout — and the only honest reply to silence is to ask again. Carefully.

"Carefully" is carrying the entire chapter on its back. Without retries, every transient blip becomes a user-visible error, and in a system where one request fans out to a dozen downstream calls, a hiccup in any one of them fails the whole thing. A payments company that treated every timeout as permanent would decline a card every time a GC pause hit a downstream service. Nobody runs a business that way. Retries are not optional.

And yet: retries amplify load, and they do it at precisely the worst moment. A service fails because it's overloaded; you respond by sending it more of the same traffic that overloaded it, multiplied by every client that just timed out. It's the infrastructure version of telling a drowning person to swim harder. The service struggling under 1,000 requests is now under 2,000 — and you haven't helped it recover. You've finished it off.

The pattern that threads the needle
BACKOFFExponential, with jitter. Spread the retries over time, and scatter them so a thousand clients don't retry as one.
BUDGETA retry budget. Cap the fraction of traffic allowed to be retries, so amplification can't run away.
DEADLINEDeadline awareness. Don't retry for your own benefit when there's no one left on the other end to benefit.

Fenced by a circuit breaker (Chapter 4) and made safe to repeat by idempotency (Chapter 2) — none of these come free. The road from "retry on timeout" to "retry without starting a stampede" runs directly through at least one production incident. Most teams pay that toll in person. This chapter is an attempt to let you skip the line.

The Problem, Actually

"Retries cause load" is true and useless. Here's the specific shape of the disaster. Service A calls Service B. B slows down — not crashes, slows, which is worse — because a slow query crept onto the hot path, or a replica fell behind. A's calls start timing out. A retries. Immediately. Because that's what the code does, and the person who wrote the code was not thinking about today.

B was handling 1,000 requests a second. Now it's those 1,000 plus 1,000 retries from the clients that just timed out — 2,000 aimed at a service already buckling at 1,000. B slows further. More clients time out. More retries. Inside a minute you've gone from "B is a little slow" to "B is gone, and the retry storm is the outage now."

Interactive · the retry-storm amplifier Baseline held at 1,000 RPS
Retries per failed request3 retries
Fraction of requests failing50%
The storm isn't caused by more users — baseline never moved. It's manufactured entirely by the retries. "Retry 3 times" reads as innocent and is a 4× multiplier on a service with no headroom to give.

The detail that makes this miserable to debug is that cause and effect come apart in the dashboards. The slow query that kicked it off may have cleared in thirty seconds. The retry storm it lit runs for minutes, because the retries are now generating the load that generates the timeouts that generate the retries. The original problem is long dead. The incident is running on its own exhaust — a perpetual-motion machine whose only output is pages. I once spent the first forty minutes of an outage hunting a downstream that had been healthy for thirty-nine of them.

The first fix anyone reaches for is to wait between retries, and it doesn't work for a reason nobody sees coming: all 1,000 clients that timed out run the same code, so they all wait the same second, then all retry at the same instant. You didn't spread the load — you gathered it up, held it for a second, and delivered it as a single synchronized punch. This is the thundering herd, and fixed delays summon it with eerie reliability.

The Naive Solutions & What They Cost

Every engineer who has implemented retries has written approximately this sequence, in this order, usually one production incident apart.

01
Just call it again.Works exactly when the problem was a single dropped packet that already healed — often enough to seduce, rare enough to trap. The moment the downstream is genuinely overloaded, the retry lands before the server finished choking on the original, and you've doubled your contribution for the price of one caught exception. (Appendix A.1)
02
Add a delay.The engineer learned "wait a bit" — real progress. What they haven't learned is that a thousand callers running this exact loop produce a thousand retries one second after the timeout, to the millisecond. The delay didn't disperse the herd. It scheduled it. (Appendix A.2)
03
Make the delay exponential.Sleep 1, 2, 4, 8, 16. Retries spread across time — the first version that could plausibly give a sick service room to stand up. The flaw is correlation: clients that failed together back off together, on the same schedule, and arrive in waves spaced just far enough apart to show up as a tidy heartbeat. When your incident's retry pattern is regular enough to set a watch by, that's the tell. (Appendix A.3)
04
Add jitter.This one works. The randomness breaks the correlation, the clients that timed out together now retry at different moments, and the load smooths into something a recovering service can survive. Its only remaining sin is the hardcoded range(5) — an attempt count connected to nothing. (Appendix A.4)
Why experienced engineers keep rediscovering this from the inside

Every step in that sequence is locally reasonable. Each client, considered alone, behaves impeccably. The catastrophe is entirely emergent — it lives in the correlation between a thousand well-mannered clients, none of which can see the other 999. You can read every retry block in the codebase, find nothing wrong with any of them, and still have a storm waiting for the next slow Tuesday. The bug isn't in the code. It's in the arithmetic of everyone running the same correct code at the same moment.

Failure Modes Worth Naming

✕Retry amplificationTHE ARITHMETIC

A hundred clients, each willing to retry ten times, turn 100 requests into 1,000 — ten times the load, aimed at a service whose entire problem was that it already had too much. A healthy-looking 20% of headroom does not survive contact with a 10× multiplier. The math is indifferent to the fact that every individual client was being sensible; sensible, multiplied by a hundred and correlated in time, is a DDoS you built yourself and now pay to host.

The deadline mismatch

Full exponential backoff over five attempts is 1 + 2 + 4 + 8 + 16 = 31 seconds of waiting — a peculiar amount of time to spend when the user's browser gave up at 30. You've built retry logic that structurally cannot succeed before the only person who cared has left the building. The retries fire anyway, because the service has no idea the caller is gone — burning threads and downstream capacity to compute an answer with nowhere to go. It's the office worker still polishing slides for a meeting that ended an hour ago.

Retrying non-idempotent operations

A POST /payments returns a 504; the original charge actually succeeded but the response evaporated on the way back; the retry charges the card again. Double charge, over-decremented inventory, a refund nobody's proud of. "Safe to repeat" is not a property you get for free — it's idempotency keys, dedup logic, and a handler that treats at-least-once delivery as exactly-once on purpose. Bolt retries onto a system that never thought about idempotency and you haven't added resilience; you've added a data-corruption bug on a timer.

The cascading retry storm

Amplification with a blast radius. A retries B; while B strains to recover, its own calls to C and D are timing out, so B retries those; meanwhile C and D have other clients also retrying. Every hop multiplies the load on the hop beneath it. The thing that takes down the company was not the service with the original hiccup — it was every service standing behind it, each faithfully amplifying the panic of the one in front.

Figure · amplification across hops is multiplicative, not additive
NO BUDGET · ×3 at every hop
A 100 ×3➔ B ~300 · falls ×3➔ C ~900 D ~900 up to 27× at the bottom
BUDGET AT EACH HOP · ×1.1
A 100 ×1.1➔ B ~110 ×1.1➔ C ~120 D ~120 exponent → rounding error
Three hops of "just retry three times" is up to 27× at the bottom. A budget at each hop is what turns that exponent back into a rounding error.
Part II · The Patterns

Pattern 1 · Exponential Backoff with Jitter

The mechanism is the one AWS wrote up in 2015, in a post that remains the clearest thing anyone has published on the subject: wrap the call in a loop, and on each failure sleep for a random interval drawn from zero up to an exponentially growing ceiling, capped so it can't run away. (Appendix A.5)

base
The floor on the first backoff — usually a second, lower for fast internal calls where a full second is an eternity.
cap
The ceiling, usually ~60 seconds, there to stop the exponential from doing what exponentials do: without it, ten doublings carry the window past seventeen minutes — a schedule nobody designed, documented, or is watching the graph for.
Interactive · what the downstream feels 1,000 timed out together
Clients that failed at the same instant1,000
All three send the same number of retries. What changes is whether they arrive as a punch or a drizzle — and only the drizzle is survivable.

The word doing the real work is full. "Full jitter" means the wait is uniformly random across the whole window — random(0, window) — not the exponential value with a little noise sprinkled on top. People conflate the two and assume the difference is cosmetic. It isn't: full jitter spreads retries far more evenly at the tail, which is exactly where a recovering service is most fragile. The gap is a few characters of code and a meaningfully gentler load curve.

Why randomness works at all

A thousand clients each making an independent random draw from the same interval will, with no coordination whatsoever, spread themselves smoothly across it. No central scheduler, no service-mesh magic, no clients talking to each other. Each one acts in pure isolated self-interest, and the aggregate is cooperative. It's one of the few places in distributed systems where selfishness composes into something polite, and it costs exactly one call to a random number generator.

Pattern 2 · Retry Budgets

The trouble with max_attempts=5 is that it's a number with no referent. Nobody computed it; somebody picked it. Retry budgets throw out the per-client count and ask a system-level question instead: what fraction of my total traffic am I willing to let be retries? Google's SRE book documents the pattern, and the recommended ceiling is 10% — with the pointed footnote that even 10% is high for a healthy system. (Appendix A.6)

Figure · the budget moves the rejection from B back to the client
NO BUDGET · retries crash through B's capacity
400 RPS to B cap 110
10% BUDGET · latches shut, holds flat under the line
110 RPS to B excess fast-failed at the client
The client swallows its own excess retries as immediate failures so B never sees them. Amplification is bounded by a ratio you chose, not by how many clients happen to be panicking.

A service taking 100 requests a second with a 10% budget sees on the order of 110 — not 400, not 1,000 — no matter how many individual clients are panicking. And the real ceiling is tighter than that clean number suggests: as the dependency sickens and its successes dry up, the retry ratio crosses the line faster, so the budget clamps down harder at exactly the moment you need it to. The budget doesn't care how many clients there are. It caps the sum.

The math is the smaller half of the value. The larger half is that the budget forces a conversation a team would otherwise never have: what fraction of our requests are we willing to spend on retries? A team that has answered that has started treating retry behavior as a property of the system. A team that hasn't will be introduced to the question in production, by the system, at a time of the system's choosing — and the system has a cruel sense of timing.

Pattern 3 · Deadline-Aware Retries

The enabling mechanism is context propagation: the original caller's deadline rides along through the entire call chain, so every service knows not just that it should hurry but exactly how much time is left. gRPC does this natively; HTTP services do it by convention — a header (X-Request-Deadline) that every hop reads, honors, and passes forward, and that exactly one team always forgets to forward. With a deadline in hand, the retry decision grows a precondition: before sleeping, check whether there's enough time left to bother. (Appendix A.7)

Figure · spending the caller's 10-second deadline
DEADLINE-BLIND · keeps firing past the line
a1
a2
a3
a4 · past 10s
result ready at 12s · caller left at 10s · answer discarded
DEADLINE-AWARE · refuses the attempt that would cross the line
a1
a2
a3
time we didn't waste
remaining < next backoff → fast DeadlineExceeded at 7s
Past a certain point an extra retry has zero chance of helping the caller and a 100% chance of costing you resources. The deadline check is how a service refuses to spend time it can't turn into a useful answer.

If you've got two seconds left and the next backoff is four, you will certainly miss the deadline — so fail now, immediately, rather than sleeping four seconds to produce an answer the caller abandoned two seconds ago. That check is a small piece of code and a genuine shift in posture: don't retry for your own benefit when there's no longer anyone on the other end to benefit. It kills a whole category of waste where services deep in the graph retry heroically for a request the API gateway gave up on ages ago. It's a kitchen still cooking an order for a table that paid and left.

Part III · Decisions

Trade-offs Worth Arguing About

Retry count vs. amplification

Take a downstream at 70% utilization — 30% headroom, which on paper reads as healthy. A client that retries once on failure sends, at the bad moment, up to 2× its normal volume. 2× into 70% is 140%, which is over the line. Read that twice, because it's the whole trap: the retries are the cause, not the response. The service didn't fail and then get retried. It got retried and then failed.

Google's published guidance, worth stealing

Three retries for most off-critical-path services; one retry for anything on the critical path. The reasoning is latency, not load — on a critical path, the time spent grinding through backoff cycles can blow past the user's patience, so you're better off failing fast with a clean error than spending ten seconds retrying your way to the same error with worse manners.

Which backoff strategy

Sounds like it should be a long debate; usually isn't. Exponential with full jitter is the right default for anything under load. Linear backoff earns its place only in the narrow case where you care about fairness over spreading — rate-limited clients with a known reset window. Pure exponential without jitter is the one to be suspicious of: it looks like the sophisticated choice and quietly keeps the clients correlated, so it helps less than its reputation promises.

Which errors to retry — the status code is a hint, not a contract

Code
Retry?
400
Almost always wrong — the request was malformed and will be exactly as malformed the second time.
404
A judgment call on your consistency model: reasonable in an eventually-consistent system where the record might not have propagated yet, pointless in a strongly-consistent one.
503
Usually worth retrying; it's practically an engraved invitation.
500
Depends entirely on whether it's a transient handler failure or a deterministic bug that will greet every identical request with identical enthusiasm.

Retrying everything uniformly is how you turn one bad request into the same bad request, sent five times. The policy has to be fitted to the actual error space of the actual service.

What the Companies Actually Built

Google SRECap amplification at the service, not the client

Let every client retry N times on its own judgment and your worst-case amplification is N×, which for a service already underwater is frequently the difference between recovering and not. The budget pulls that ceiling down to a number the service itself controls and the clients cannot override by panicking harder. The documented figure is 10%, and in practice a healthy service runs well below it — a sustained retry rate above 1–2% is already worth a look.

The tell is the retry success rate

If your retries are mostly succeeding, the dependency is flaky and the retries are doing their job. If retries fail at the same rate as first attempts, you're not retrying — you're generating load with extra steps, and you should be failing fast instead.

The part that's organizational rather than technical — and therefore the part that doesn't get done — is that budgets need client and server teams to talk. The client owns the retry behavior; the server owns capacity and has to expose the retry rate as a metric. When they operate in silos, they meet for the first time during the incident, which is the most expensive venue ever devised for an introduction.

StripeIdempotency first, then retry

Stripe starts one place earlier than most teams: idempotency is a prerequisite, not a follow-up ticket. Every charge carries a client-generated idempotency key, and the server deduplicates on it for a configured window, so retrying a payment is provably safe — a duplicate gets recognized and the original result replayed, rather than charging the customer a second time for the privilege of a flaky network. On that foundation they layer deadline propagation: a payment with a 10-second timeout won't be retried at 8 seconds if the next backoff is 4.

The transferable lesson: orderingBuild idempotency first, then retries — never the reverse. Bolt idempotency onto a system that already retries and you're auditing every handler for hidden side effects, one grim spreadsheet row at a time. Get the contract right and retries stop being dangerous and start being boring, which is the highest compliment you can pay a distributed-systems primitive.
NetflixAdaptive backoff — a feedback loop, not a schedule

Netflix's retry behavior is less a fixed schedule and more a feedback loop. Their Hystrix library — since largely handed off to Resilience4j — tracked per-dependency health and latency in real time and used those signals to decide how hard to keep leaning on a struggling dependency. As error rate and latency climbed, the breaker tightened and clients eased off; as it recovered, they leaned back in, without a human in the loop. The intuition is that a fixed backoff treats a healthy service and a dying one identically, which is faintly absurd when you say it out loud — distinguishing the two is the entire job.

The cost is complexity, and it's not a rounding errorLatency telemetry per dependency, per-service tuning, and enough monitoring to notice when the adaptive machinery itself misbehaves. Netflix can sign that check because their dependency graph is enormous. At most companies' scale, fixed exponential backoff with jitter is the right answer, and adaptive backoff is a beautifully engineered solution to a problem you're allowed to not have yet. Knowing which of those you are is the actual skill.
Part IV · Operating It

The Principal Engineer's View

The question that matters most is not "which backoff algorithm should we use?" It's "does anyone here actually know what our retry behavior is doing to our dependencies right now?" At most companies the honest answer is no — retry behavior is buried three layers deep in client libraries, set to a default sensible for the service it shipped with and arbitrary for yours, and surfaced on precisely zero dashboards. The first time anyone looks at it directly is mid-cascade, reverse-engineering it from the wreckage.

Three metrics that deserve the observability you'd never skip for latency

RATERetry rate, per client, per dependency. What fraction of outbound calls are retries? This is the number that's invisible today and obvious in hindsight.
SUCCRetry success rate. Of the retries you send, how many succeed? A low success rate next to a high retry rate is the signature of a storm: you're spending load to accomplish nothing.
LOADRetry contribution to downstream load. 1,000 RPS at a 15% retry rate is 150 RPS of retries your own error rate never predicted — and somebody should know that number before it matters.
The red flag is the specific combination

High retry rate, low retry success rate. That isn't resilience working. That's a storm in progress, and the retries aren't helping — they may be the only reason it isn't recovering.

Questions to put to your team, out loud, before an incident does

01What's our retry budget? Has anyone multiplied our retry rate against downstream capacity to see what it does under load?
02Are we retrying idempotently? Does the team that owns the downstream even know we retry, and did they build for it?
03Do our retries respect caller deadlines, or are we burning compute on answers no one is still waiting for?
04How would we detect a retry storm? What alert fires, and who does it wake?
05Have we tested any of this under real load? When a dependency degrades in a chaos experiment, what does our retry behavior actually do — not what do we assume it does?

The last one is the question most teams can't answer with a straight face, because retry behavior is almost always validated in staging, where the load is a rounding error and the thundering herd can't form for lack of a herd. You need enough concurrent clients to stampede before you can watch a stampede, and you only have those in production. Which is the entire case for chaos engineering — because the alternative is letting the system pick the moment, and the system always picks 2 a.m.

Part V · Practice

Exercise · Keep Within the Budget

SCENARIOService A → Service B

Service A sends 100 RPS to Service B. B can sustainably handle 110 — call it 10% headroom. B has a brief degradation. A's current policy: up to 3 retries on any 5xx, fixed 1-second delay between attempts.

WORK ITIf every request fails and triggers all 3 retries, that's 100 original + 300 retries = 400 RPS into a service that tops out at 110. The storm pushes B to nearly 4× capacity — it does not recover, it cascades, and now the retry policy is the outage.
YOUR TASKNever exceed 110 RPS to B, under any failure you can construct
HINTA 10% retry budget holds retries to roughly a tenth of traffic — on the order of 110 RPS total — and clamps harder as B's successes dry up. Layer full jitter on top and even those retries spread across time, so the instantaneous peak sits below the average. Then add the circuit breaker from Chapter 4: once B is clearly failing rather than flickering, stop retrying altogether and fail fast.

The point isn't the answer — it's the argument you have with yourself getting there. Paste it into an AI and you'll get a clean, confident response that skips the only thing that matters: your dependencies, your traffic shape, your blast radius.

Connections to Other Chapters

← CH 2Idempotency. Retries are safe only when the operation is safe to repeat. Before you add a retry to any call, confirm the handler on the other end was built for at-least-once delivery. If it wasn't, the retry isn't resilience — it's a data-corruption bug on a delay timer that goes off the first time a response gets lost in flight.
← CH 4Circuit Breakers. The breaker is what stops the retries when a service is past saving. Retries without a breaker are how you get the infinite storm; the breaker is the off switch. Backoff governs how aggressively you retry while there's still hope, and the breaker decides when hope has resolved into "just stop."
→ CH 12Dead Letter Queues. When the retries are exhausted — budget spent, circuit open, operation still not done — the work has to go somewhere. A dead letter queue is the somewhere: park the failed operation, let the dependency recover on its own schedule, and process it later instead of blocking a caller on it now.
The intuition to walk away with

A retry is never a private decision. It feels like one — a few lines at a single call site, handling a single failure — and that's exactly the illusion that keeps retry storms perennial. The blast radius of those few lines is every service downstream of you, multiplied by every other client running their own reasonable version of the same few lines, correlated by the shared instant you all failed together. Backoff, jitter, budgets, and deadlines aren't four tricks; they're four ways of forcing a local decision to account for a global consequence. Get them right and a thousand clients failing at once produce a survivable drizzle. Get them wrong and they produce a service that was fine until its friends tried to help.

Appendix A · Reference Implementations

The chapter keeps code out of the narrative on purpose. Here it is — collected, runnable, annotated. All of it is illustrative rather than production-hardened: the snippets skip the timeouts, error taxonomy, thread safety, and metrics you'll need in the real thing, and the first four are deliberately broken in instructive ways.

A.1Immediate Retry · the optimist's reflexPython

Works only when the failure was a single self-healing packet drop; the instant the downstream is actually overloaded, the retry arrives before the server finished the original and you've doubled your load for the price of one caught exception.

A.2Fixed Delay · the herd, rescheduledPython

The pause feels responsible, but a thousand callers running this loop all wake one second after the timeout and retry in unison. The delay preserves the thundering herd intact — it just slides it one second to the right.

A.3Exponential Backoff, No Jitter · correlated wavesPython

Finally spreads retries across time, but clients that failed together share one schedule and arrive in synchronized spikes at t+1, t+3, t+7. The giveaway is a retry pattern regular enough to set a clock by.

A.4Exponential Backoff with Jitter · the first one that worksPython

The random offset breaks the correlation, so clients that timed out together retry at different moments and the load smooths out. The version to actually ship — its only remaining sin is the hardcoded range(5), an attempt count tied to nothing real.

A.5Full Jitter · the canonical formPython

The AWS-2015 formulation: wait a uniform-random interval across the whole exponential window — random(0, window), not the window plus noise. base floors the first backoff; cap stops ten doublings from quietly scheduling a seventeen-minute wait.

A.6Retry Budget · amplification, capped at the sourcePython

Reframes retries from a per-client count into a system-level rate: spend retries only while they're under some fraction of total traffic (Google uses 10%), and fail fast once the budget's gone. This toy version skips the sliding window, thread safety, and per-dependency tracking the real thing needs.

A.7Deadline-Aware Retry · check the clock before you waitPython

Before sleeping, check whether enough of the caller's deadline remains to make the attempt worth it; if not, fail immediately instead of computing an answer no one is waiting for. Requires the deadline propagated down the chain — gRPC does it natively, HTTP by a header someone always forgets to forward.

Next: Chapter 6 — Request Coalescing: turning a thousand identical in-flight requests into one piece of work, shared.

Previous chapter
←  04 · Circuit Breakers
Next chapter
06 · Request Coalescing →