Reliability

Retry storms: when recovery makes the outage worse

One customer action can turn into dozens of backend attempts. Count them before deciding the system needs to try harder.

7 min read

The practical answer

Retry storms happen when repeated attempts add enough work to prolong or worsen a failure. Control them with an explicit retry owner, bounded attempts, a total deadline, and a delay policy with jitter. Verify that repeated writes are safe and measure successful operations alongside attempt volume and queue growth.

I’m persistent. If something doesn’t work, I want to understand it and try again. That mindset has carried me through years of failed projects and things I had to teach myself. It also makes me pay attention to what changes between attempts. Effort has to produce useful information or useful work.

A retry loop doesn’t make that judgment for us. It follows the rules we give it, including rules scattered across clients, services, and workers. When those rules send more work into a struggling dependency, retry storms can turn a recoverable problem into a longer outage. I want the retry policy to be something the team can explain from the customer’s action all the way down.

1. Count how retry storms multiply the work

Start with a single logical operation, such as generating one report. Find every layer that can repeat part of it: the browser client, API service, dependency SDK, and job runner. Record whether each setting means retries after the first call or total attempts including it. That small wording difference matters.

Hypothetical worst case: three nested layers, three total attempts each
LayerBehavior when every attempt failsCumulative calls
ClientInvokes the API up to three times.3 API calls
APIInvokes a worker service up to three times per API call.9 worker-service calls
Worker serviceInvokes the dependency up to three times per worker call.27 dependency calls

That arithmetic assumes independent retry loops, every attempt failing, and no shared deadline or quota stopping the chain. It describes a possible maximum for this example. It does not mean every request in a three-service application produces 27 calls. Inspect your actual call path and settings.

Google’s overload guidance describes this multiplication problem and uses retry budgets to limit extra work. The lesson I take from it is to make ownership visible. If several layers can retry the same downstream failure, decide which layer should do so and how the others learn that the budget is exhausted.

Reference: Google SRE: Handling Overload

2. Give the operation a budget it cannot quietly reset

I want a retry policy written in terms someone can review: which failures qualify, how many total attempts are allowed, how much elapsed time is available, and which component enforces those limits. Check the deployed library configuration. A wrapper with two retries may sit around an SDK that already retries internally.

AWS SDK documentation describes error classification, attempt limits, backoff, and retry quotas. Behavior depends on the SDK and its configuration, so read the documentation for the version and mode you use. A retryable transient failure needs a different response from a request that will keep failing validation unchanged.

Reference: AWS SDKs: Retry behavior and configuration

For a hypothetical operation with a two-second caller deadline, a retry that starts after 1.9 seconds has very little time left to help. Carry the remaining budget into the next call and include connection setup, waiting, and backoff. If there isn’t enough time for a useful attempt, finish with the defined failure or deferred-work response.

Exponential backoff spaces attempts further apart, while jitter varies their timing so clients are less likely to retry together. Cap the delay and the number of attempts. AWS’s Builders’ Library explains why synchronized retries can keep pressure on an overloaded dependency even when each client has a delay.

Reference: AWS Builders’ Library: Timeouts, retries, and backoff with jitter

Also check what cancellation actually stops. Ending the caller’s wait may leave downstream work running. That behavior belongs in the budget discussion, especially if another attempt can overlap it.

3. Find out what a timed-out write may already have done

A timeout tells the caller it didn’t receive a result in time. The server may already have performed the action. AWS’s retry guidance calls out this uncertainty because repeating an operation with side effects can create an unwanted second result.

Reference: AWS Builders’ Library: Timeouts, retries, and backoff with jitter

Imagine a hypothetical create-export request. The server stores the export job, but the response is lost. A second request needs a way to identify the same intended operation. Generating a fresh identifier for every network attempt would lose that connection. Keep the logical operation’s identifier stable across its retries.

AWS’s idempotent API guidance uses a caller-provided request identifier scoped to the caller. It also explains why recording that identifier and performing the associated mutation need an atomic boundary, why reusing a key with different parameters should fail validation, and why the retention period needs an explicit contract.

Reference: AWS Builders’ Library: Making retries safe with idempotent APIs

For our export example, I’d review concurrent requests with the same key, a reused key with different options, a lost response, and a request arriving after the key expires. Then inspect the worker separately. Deduplicating job creation doesn’t automatically make every action inside the job safe to repeat.

If the workflow includes another service, identify the boundary each idempotency guarantee covers. Define how to look up an uncertain result or reconcile it before adding more attempts. A header named Idempotency-Key only helps when the receiving service implements the promised behavior.

4. Leave enough capacity for recovery to work

Backoff changes timing. It doesn’t tell you whether the system has enough capacity to handle all the pending work. Watch the incoming logical operation rate, the attempt rate, and the age of queued work together. If the queue keeps growing, a slower retry loop may still be accumulating a problem.

Google’s overload guidance covers per-request and per-client retry budgets, as well as rejecting work when a service cannot accept more. Those controls address different scopes: a small attempt limit on every request can still add substantial load across many clients.

Reference: Google SRE: Handling Overload

For a hypothetical reporting service, decide how many exports may run at once, how many may wait, and how long an accepted export remains useful. Make the response match reality. Telling a user their export is accepted creates an obligation to track it, finish it, or explain its failure. Letting a queue grow without a limit postpones that decision.

During recovery, resume deferred work at a rate the dependency can sustain. Test the recovery phase as deliberately as the failure phase. Releasing the entire backlog as soon as a health check turns green can put the same pressure back on the system. I want the plan to cover what happens after the first successful request.

5. Verify customer results through failure and recovery

Use a controlled environment to introduce a slow response, a transient failure, and a response lost after a successful write. Trace one logical operation through every attempt. These cases ask different questions, so I would keep their expected results explicit.

  • Does the attempt count stay within the configured budget across layers?
  • Does the operation stop scheduling new attempts when its deadline expires?
  • Can an uncertain write be resolved without creating an unwanted duplicate?
  • Does pending work remain within the queue and concurrency limits?
  • When the dependency recovers, does the backlog drain while useful work completes?

Compare successful customer operations and terminal failures alongside latency and attempts per operation. Fewer retries alone could mean the system has stopped trying too soon. More successful attempts could still conceal duplicate jobs. The measurements should tell you whether the customer’s intended action finished correctly.

A retry policy the whole team can inspect
Logical operation and dependency:
Retry owner and retries inside libraries:
Eligible failures:
Maximum total attempts, including the first:
Overall deadline and per-attempt limits:
Backoff, jitter, and shared retry quota:
Idempotency scope, retention, and uncertain-result lookup:
Queue and concurrency limits:
Failure and recovery test results:

I’m all for persistence. In a backend, that means giving each additional attempt a reason to exist and a limit it has to respect. Once those rules are visible, the team can test them, explain them, and change them with evidence.

References and further reading

All guides