Hire Software Developers 7
Back to blogs

Timeouts, retries and fallbacks: stopping one dependency taking you down

A single orange block wedged in a narrow gate between two pale pillars, with a large crowd of blue blocks packed in behind it, representing retries piling up against one slow dependency

Timeouts, retries and fallbacks: stopping one dependency taking you down

At 14:02 the payment provider's API slows down for about ten seconds. By 14:03 it's answering normally again. At 14:40 your checkout is still throwing errors, the provider's status page is green, and someone in the incident channel types the sentence nobody wants to see there: "but their side recovered forty minutes ago."

The postmortem will blame the provider. It shouldn't: what kept you down for forty minutes was your own timeouts and retries, working exactly as configured. The provider was only the trigger, and the people who study this failure shape say blaming the trigger is the common first reading and the wrong one: "It is common for an outage that involves a metastable failure to be initially blamed on the trigger, but the true root cause is the sustaining effect." [R23]

A cascading failure is "a failure that grows over time as a result of positive feedback." [R2] A metastable failure is the version that stays broken after the cause has gone: a trigger pushes the system into a state where "a sustaining effect ... prevents the system from leaving the bad state." [R22] Timeouts, retries and fallbacks are the three settings that decide which one you get when a dependency slows down.

Key takeaways

  • In the only systematic study of these failures I found, 11 of the 21 metastable incidents Huang et al. catalogued from public incident reports were sustained by retries. That's a share of incidents the researchers selected, not an outage rate, and no published outage rate turned up. [R36] [R37]
  • Retries multiply across layers: four attempts at each of three layers is 64 attempts on the database in Google's example [R7], and "three retries at each layer" of a five-deep stack is 243x the load in AWS's [R49]. Retry at one layer.
  • Defaults won't save you. Python Requests never times out unless told to, gRPC sets no deadline, and .NET's HttpClient waits 100 seconds. [R60] [R58] [R61]
  • In Google's own arithmetic, adding a per-client retry budget of 10% to a three-attempt cap cuts retry growth from "just below 3X" in its worst-case example to "just 1.1x in the general case". [R18]
  • The evidence base is vendor practice guidance, one peer-reviewed vision paper and one small peer-reviewed study. No source measures how much any of these mitigations reduces outages or recovery time.

How timeouts and retries keep an outage alive after its cause

Bronson and colleagues, writing from years of operating distributed systems at scale (Bronson was formerly at Facebook), give the cleanest worked example of timeouts and retries turning on you. [R21] A database answers in under 100 ms below 300 queries per second and gets an order of magnitude slower above it. Each user request makes one query, plus one retry if the first hasn't returned in a second. The app runs happily at 280 QPS. Then a network switch drops out for ten seconds. When it comes back, the backlog of requests and retries hits the database at once, latency climbs past the one-second timeout, and "client queries will continue at 560 QPS due to retries." [R25]

The switch is fine now. The database is fine too, as hardware. The system stays down anyway, because every query times out and gets retried, and the retries keep latency above the timeout. By the paper's arithmetic the system was only safe below 150 QPS, and getting out requires "reducing the web application load to under 150 QPS or limiting retries to less than 20 QPS." [R25]

Read that again with your own dashboard in mind. At 280 QPS nothing was wrong. No alert, no errors. The system was healthy and one blip away from a state it couldn't leave on its own. Bronson's group calls this the vulnerable state, and they're blunt about why teams live in it: "many production systems choose to run in the vulnerable state all the time because it has much higher efficiency than the stable state." [R24]

That's a retry storm seen from the inside. Google's SRE book tells the same story with its own numbers. A backend accepts 10,000 QPS per task. The frontend sends 10,100. The extra 100 are rejected and retried, then 200, then 300: "the point remains that retries can destabilize a system." [R4] Turning the traffic down doesn't necessarily fix it: "Even if the rate of calls to MakeRequest decreases to pre-meltdown levels (9,000 QPS, for example) ... the problem might not go away." If the backend is spending its capacity on requests that will fail anyway, or retries are amplifying an already unstable backend, then "you must dramatically reduce or eliminate the load on the frontends until the retries stop and the backends stabilize." [R5]

The most unsettling number comes from a controlled experiment. Huang et al. ran a replicated database at about 6,200 successful requests per second, with clients on a 3-second timeout retrying up to 4 times. [R40] Cutting CPU by 78% for 10 seconds caused a dip and a recovery. Cutting it by 80% for the same 10 seconds didn't recover at all: latency sat at the 3-second timeout, attempted requests peaked around 20,000 RPS, and "goodput is reduced by ~90% to 600 RPS." A 9-second trigger at 80% recovered. In the authors' words, "a 2%-decrease in available CPU or a 1-second increase in duration separated successful recovery from a metastable failure." [R40]

In that experiment, a two-point difference in the CPU cut separated a blip from an outage. That's why the forty-minute checkout incident feels unfair. The trigger doesn't need to be much bigger than last month's blip. Slightly bigger is enough.

One more line from Google, because it explains why these incidents run long: "Graphs of retry rates can be an indication of bad retry behavior, but may be confused as a symptom instead of a compounding cause." [R9] The retry graph is right there on the dashboard, and it's easy to read as a consequence of the problem rather than part of it.

What the evidence is, and what nobody has measured

Before any advice, the honest inventory, because this topic runs on confident numbers with thin provenance.

The Google SRE book is practice guidance from Google's own engineers, not research. Chapter 22, "Addressing Cascading Failures", was written by Mike Ulrich, and the book is published by O'Reilly and free online under a Creative Commons licence. [R1] Its numbers are worked examples. The retry cascade comes with the note "Some simplifying assumptions were made here to illustrate this scenario," and a footnote that reads: "An instructive exercise, left for the reader: write a simple simulator." [R17] Authoritative about how Google runs Google. Not a measurement of anyone else's systems.

AWS's Builders' Library is the same kind of source. Marc Brooker's "Timeouts, retries, and backoff with jitter" (2019) is AWS describing AWS practice. [R44] Valuable, specific, and not independent.

Bronson et al. is a vision paper. It appeared at the HotOS '21 workshop, runs to seven pages, and draws on "examples observed during years of operating distributed systems at scale." [R21] It discloses no sample. The follow-up study says so directly: Bronson et al. "only asserted that the pattern was common, but did not present any data about real-world occurrences." [R35] Bronson's own paper describes the existing practitioner discussion as "SRE folklore" and notes that "the study of these large-scale failures have so far eluded academia." [R28]

Huang et al. is the only systematic look I found, and it's small. Published at OSDI '22 by researchers from Penn State, the University of New Hampshire and Twitter [R33], it reports "22 metastable failures from 11 different organizations" and states that "at least 4 out of 15 major outages in the last decade at Amazon Web Services were caused by metastable failures." [R34] That line appears in the abstract; the paper doesn't list the 15, and it covers metastable failures of any mechanism, not retries specifically. The method matters. The authors worked "by sifting through hundreds of publicly available incident reports", identified 21 incidents for their main table (some flagged only as a "plausible metastable failure"), and state that "we use our best judgment in identifying metastable failures." [R36] Of those 21, 11 were sustained by retries, which the paper summarises as "more than 50% of the studied incidents." [R37] The outages in their set ran from 1.5 to 73.53 hours. [R38]

Now what those numbers are not. Eleven of twenty-one is the share of hand-picked metastable incidents in which retries were the sustaining effect. It isn't the share of outages caused by retries, because the denominator is incidents the researchers had already classified as metastable, drawn only from organisations that publish incident reports. I found no published figure for how often retries cause outages, and none of these sources gives a figure for how much timeouts, retry budgets, jitter or circuit breakers reduce outage frequency or recovery time. If you've seen one, check where it came from.

So you're choosing mechanisms without an expected-value calculation. That's fine. The mechanisms are well understood even where the frequencies aren't, and the next three sections are about mechanism.

Timeouts: your client's default is probably "wait"

Start with what you get for free, because it's less than you'd hope.

Python's Requests library: "If no timeout is specified explicitly, requests do not time out." Its own documentation warns that "Nearly all production code should use this parameter in nearly all requests. Failure to do so can cause your program to hang indefinitely." [R60] gRPC: "By default, gRPC does not set a deadline which means it is possible for a client to end up waiting for a response effectively forever." [R58] .NET's HttpClient does set one, and it's 100 seconds: "The default value is 100,000 milliseconds (100 seconds)." [R61]

A hundred seconds sounds like a safety net. Google's arithmetic shows why it's closer to a trap. Take a frontend with 1,000 worker threads, serving 1,000 QPS at 100 ms each, so about 100 threads are busy. Now let 5% of requests hang because one slice of a backend is unavailable. With a 100-second deadline, those 5% want 5,000 threads. The frontend has 1,000. "Assuming no other secondary effects," it can serve 19.6% of requests, "resulting in an 80.4% error rate." [R12] A problem affecting one request in twenty becomes an error for four in five, because the timeout let the broken requests hold every thread. Google's rule of thumb: "Having deadlines several orders of magnitude longer than the mean request latency is usually bad." [R12]

The wasted work is the other half: "you don't get credit for late assignments with RPCs." [R10] A server finishing a request whose caller gave up long ago is burning capacity on an answer nobody will read, during exactly the minutes it needs that capacity.

How long should a timeout be?

Set it from the dependency's measured latency, not from a round number. AWS's method: "we choose an acceptable rate of false timeouts (such as 0.1%). Then, we look at the corresponding latency percentile on the downstream service (p99.9 in this example)." [R46] The same article lists where that breaks: calls over the public internet, services where "p99.9 is close to p50", and timeouts that don't cover everything, "like DNS or TLS handshakes." [R46]

Brooker's own war story is the TLS one. A dependency timeout of about 20 milliseconds that only fired after deployments turned out to include "establishing a new secure connection, which was reused on subsequent requests." The first workaround was a bigger timeout. The real fix was "establishing these connections when a process started up, but before receiving traffic." [R47] The timeout was measuring something other than what its owner thought.

Two more rules from the SRE chapter. Propagate the deadline rather than inventing one at each hop: if server A sets 30 seconds and spends 7 before calling B, B gets 23, and if B spends 4 before calling C, C gets 19. [R11] Java and Go gRPC do this by default; C++ needs it switched on. [R58] And when a backend is known to be down, don't wait for the timeout at all: "it's usually best to immediately return an error for that backend ... If your RPC layer supports a fail-fast option, use it." [R13]

How many times should you retry?

Fewer times than feels safe, and at one layer only. Google caps each request at three attempts and each client at a 10% retry ratio [R18]; AWS, "for low-cost control-plane and data-plane operations," retries "at a single point in the stack." [R49] None of these sources offers a universally right number.

The multiplication is what kills you. Google's example: if "the backend, frontend, and JavaScript layers all issue 3 retries (4 attempts), then a single user action may create 64 attempts (4^3) on the database." [R7] AWS's version is a five-deep stack with "three retries at each layer": "the load on the database will increase 243x, making it unlikely to ever recover." [R49] (AWS's arithmetic, 3^5, counts three tries per layer; three retries on top of a first attempt would be worse. The point survives either way.) Google's rule is that "requests should only be retried at the layer immediately above the layer that is rejecting them," with an explicit "overloaded; don't retry" error so the layers above give up instead of joining in. [R19]

A per-request cap isn't enough on its own. Google's arithmetic for its worst case, a datacenter rejecting a large portion of what it receives, has a three-attempt cap growing load "to somewhere just below 3X." Add a per-client budget, where a client only retries while retries are under 10% of its requests, and the growth drops "to just 1.1x in the general case." [R18] That's Google modelling its own clients, not a benchmark, but the logic carries: a budget caps retries in aggregate, where a cap per request only limits each request. Chapter 22 gives a cruder version, "only allow 60 retries per minute in a process" [R6], and gRPC ships one as configuration: with maxTokens: 10 and tokenRatio: 0.1, failures cost a token, successes earn back a tenth of one, and "If the token_count falls below half of maxTokens, retries are paused until the count recovers." [R59] Huang et al. put the bound plainly: "a policy with at most two retries will not amplify the work more than three times, while the policy with no cap effectively leaves the system with no stable region." [R39]

Then exponential backoff with jitter. The SRE book's instruction is "Always use randomized exponential backoff when scheduling retries," and for the jitter it points readers to Brooker's 2015 AWS Architecture Blog post. [R6] That post is the SRE book's cited evidence for jitter, and it's a simulation of N clients contending for one database row under optimistic concurrency, not a production measurement. [R54] In that simulation, plain exponential backoff "helps a small amount, but doesn't solve the problem," and with 100 contending clients, adding jitter "reduced our call count by more than half." He's careful about the limit, too: "none of these approaches fundamentally change the N2 nature of the work to be done." [R54] In that simulation jitter cut the work substantially. It didn't change how the work grows as more clients pile in.

Two decisions sit underneath all of this. First, which failures to retry: "Don't retry permanent errors or malformed requests in a client, because neither will ever succeed." [R8] Second, whether the call is safe to repeat. AWS's position is that "APIs with side effects aren't safe to retry unless they provide idempotency" [R51], which is the same constraint that shows up in the transactional outbox: at-least-once delivery plus an idempotent consumer. If your retry wraps a database transaction, it's the same discipline as re-running a serialisation failure in the double-booking fix: the whole unit of work, or nothing.

The cautionary tales in Huang's dataset are about fixes. After an AWS SimpleDB incident, "engineers decided that servers must continue to retry the locking service instead of giving up." The paper notes that unlimited retries "make the sustaining effect more severe," and "A similar incident (AWS3) happened to the DynamoDB database about a year later." [R41] At Spotify, engineers added logging to the error path after one incident; in the next, "the additional logging after a load spike and initial retries increased the cost of each retry." [R42] Both were reasonable responses to a postmortem. Both fed the next incident.

Fallbacks and circuit breakers: decide what "degraded" means in daylight

A timeout says when to stop waiting. A fallback says what to do instead. That second decision is a product decision dressed as an engineering one, and if nobody makes it in daylight, someone makes it at 2am.

Martin Fowler's framing is still the clearest: when the call fails, "Does it fail the operation you're carrying out, or are there workarounds you can do? A credit card authorization could be put on a queue to deal with later, failure to get some data may be mitigated by showing some stale data that's good enough to display." [R57] Queue it, serve stale, hide the widget, or fail the request. Each one is right for some dependency and wrong for another, and only someone who knows the business can pick.

The circuit breaker is the mechanism that decides when to switch to the fallback. Once failures cross a threshold, "the circuit breaker trips, and all further calls to the circuit breaker return with an error, without the protected call being made at all," and a half-open state later lets one trial call through "to see if the problem is fixed." [R56] On attribution, since it gets garbled: Fowler writes that Michael Nygard "popularized the Circuit Breaker pattern" in Release It! [R55], published in 2007 [R30]. Netflix's Hystrix appears in Fowler's piece under further reading, as an open-source implementation, not as the origin. [R55]

The SRE chapter adds the warning that should shape how you build any fallback: "Remember that the code path you never use is the code path that (often) doesn't work." [R15] A fallback that has never run in production is a hypothesis. Google's suggestion is to keep exercising it, by running a small subset of servers near overload. [R15] The cheaper version for everyone else is a test where the dependency doesn't fail fast but simply never answers. The chapter calls it blackholing, and it's specific about the risk: "Backends advertised as noncritical can still cause problems on frontends when requests have long deadlines." [R14] The recommendation widget that hangs for 100 seconds takes the checkout page down with it.

The case against: each of these has a cost

It would be easy to read the last three sections as "add timeouts, retry budgets and circuit breakers everywhere." The sources don't say that, and some argue against parts of it.

Retries earn their keep. AWS: "retries are a powerful mechanism for providing high availability in the face of transient and random errors." [R53] Google goes further for the everyday case: when only a few backend tasks are overloaded, "the preferred response is to retry the request immediately," because retries landing on different tasks become "a form of organic load balancing." [R20] Strip retries out to avoid storms and transient blips become user-visible errors.

AWS is wary of circuit breakers. Brooker writes that circuit breakers "are widely promoted to solve this problem. Unfortunately, circuit breakers introduce modal behavior into systems that can be difficult to test, and can introduce significant addition time to recovery." AWS's alternative is "limiting retries locally using a token bucket," built into the AWS SDK since 2016. [R50] A breaker is a second system with its own states, thresholds and bugs, sitting in the path of every call.

Jitter and backoff don't remove the load. Brooker's own simulation shows the work stays quadratic in the number of contending clients. [R54] Backoff buys time; the budget is what bounds the total.

The reliability features can cause the outage. Google's closing remarks list "Retrying on failures, shifting load around from unhealthy servers, killing unhealthy servers, adding caches" as changes that "might be implemented to improve the normal case, but can improve the chance of causing a large-scale failure." [R16] Bronson's worst example is a geo-distributed system that "added additional retries and failover destinations to improve its steady-state reliability" and ended up with "a worst-case work amplification of over 100×." [R27]

And you can't test your way to certainty. Metastable failures "are an emergent behavior rather than a logic bug--one cannot write a unit or integration test to trigger them" [R28], and "small scale tests don't provide much confidence that a problem cannot appear at full scale." [R31] One of Bronson's cases "defied explanation for more than two years, causing multiple outages," and the eventual fix "was a single line that changed the connection pool's policy." [R29]

Where does that leave a team with limited time? My reading, not a finding in any of these papers: the cheap, low-risk changes are the ones that bound amplification. Deadlines, one retry layer, a budget, and a "don't retry" signal. Those reduce the work the system can generate against itself without adding a new state machine. A circuit breaker is a decision worth making per dependency, after those, and with AWS's objection in view.

Where to start: one dependency, one list

You don't need a resilience programme. You need one dependency done properly, then the next. Pick the one whose slowdown costs you most (the payment provider, the auth service, the primary database) and work through it.

  1. Write down the chain. For every hop between the user and that dependency: connect timeout, request timeout, retry count, and which layer retries. If three layers retry, you've found your 64. [R7]
  2. Replace defaults with numbers. Set the timeout from the dependency's p99.9 at an acceptable false-timeout rate [R46], and check it covers DNS and the TLS handshake. [R46] [R47]
  3. Propagate the deadline so the downstream call inherits what's left instead of starting a fresh clock. [R11]
  4. Retry at one layer, capped, with jitter, inside a budget. [R6] [R18] [R19] [R59]
  5. Sort errors into retryable and permanent, and return an explicit overloaded status that tells callers not to retry. [R8] [R19]
  6. Make side-effecting calls idempotent before you let anything retry them. [R51]
  7. Pick the fallback and run it. Queue, stale, hide, or fail, chosen by someone who knows the business. [R57] Then blackhole the dependency in a test environment and watch what the caller does. [R14]
  8. Put the retry rate on the same dashboard as the error rate, so nobody mistakes it for a symptom at 2am. [R9] What else belongs on that dashboard is a separate argument.

Last, write the settings down somewhere other than one engineer's memory. Retry configuration scattered across three services and a client library is exactly the kind of knowledge that leaves with a single person.

One real task

Adding a propagated deadline, a single capped retry layer with jitter and a budget, and a tested fallback to one critical dependency is a scoped piece of work with a clear definition of done, and it's easy to keep deferring until after the postmortem. That is the shape Dev On Demand is built for: one dedicated AI-augmented engineer for $3,495/mo (two in parallel for $6,795/mo), a 3-day task cycle with task-by-task approval, and cancel any time. You can start with one real task and judge the engineer before you subscribe.

Whatever you decide, open the client configuration for your most important dependency today and find the timeout. If the answer is "the default", look up what the default is.

Frequently asked questions

What is a retry storm?

A retry storm is when clients retrying failed requests generate enough extra load to keep the dependency failing, so the retries sustain the outage they were meant to ride out. Google's SRE book shows a backend overloaded by 100 QPS whose retries grow to 200, then 300 QPS [R4], and notes that lowering traffic back to normal "might not" end it. [R5] Bronson et al. write that "One of the most common failure-sustaining mechanisms is request retries." [R26]

How long should a timeout be?

Long enough to cover the dependency's normal slow tail, and no longer. AWS picks an acceptable false-timeout rate, such as 0.1%, and sets the timeout at the matching latency percentile, p99.9 in that example. [R46] Google warns that "Having deadlines several orders of magnitude longer than the mean request latency is usually bad." [R12] Check your client's default first: Python Requests never times out by default [R60], gRPC sets no deadline [R58], and .NET's HttpClient waits 100 seconds. [R61]

How many times should a client retry?

At one layer only, with a small cap and a budget. No source gives a universal number. Google uses a per-request cap of three attempts plus a per-client limit that stops retrying once retries reach 10% of requests, which in Google's own arithmetic cuts load growth from just under 3X in its worst-case example to 1.1x in the general case. [R18] Huang et al. note that "a policy with at most two retries will not amplify the work more than three times, while the policy with no cap effectively leaves the system with no stable region." [R39]

Do circuit breakers prevent cascading failures?

They can stop a caller from hammering a failing dependency, but they aren't free, and AWS argues against relying on them. A breaker trips after a failure threshold and fails calls immediately without making them. [R56] AWS argues that circuit breakers "introduce modal behavior into systems that can be difficult to test, and can introduce significant addition time to recovery," and prefers a token bucket that limits retries locally. [R50] The pattern was popularised by Michael Nygard's Release It!, not invented at Netflix. [R55] [R30]

How often do retries cause outages?

I found no published figure. The closest evidence is Huang et al. (OSDI '22), who selected 21 metastable incidents from hundreds of public incident reports and found 11 were sustained by retries. [R36] [R37] That's a share of incidents the researchers chose, classified with their "best judgment", not a rate across outages. The same paper's abstract reports "at least 4 out of 15 major outages in the last decade at Amazon Web Services were caused by metastable failures", which covers metastable failures of any mechanism, not only retries. [R34]

Sources

back to top

Related Articles

Book 30 min with Albert
Smiling man with short dark hair and glasses wearing a black suit, white shirt, and black tie against blue background.
Tell Albert what you're shipping.
He'll read this before joining the call. Phone number comes next, on the calendar step.
↳ info@you-source.com
↳ 4-hour response
Please wait while we retrieve meeting schedules.
Oops! There's a problem with your request. We're working on fixing it. Please try again later.