The outage is over. The load balancer is green. Ten thousand clients each saw a timeout, each applied the same “wait one second,” and each fires at t=1s. The box that just recovered takes 10k retries in one instant and falls over again. Nobody chose a DDoS. They synchronized the wake-ups.
Wait min(cap, base × 2^attempt) after a failure, then add jitter so the same attempt does not fire at one instant. Double the ceiling on each miss. Cap it so attempt 20 is not a day-long nap. Draw a random wait inside that ceiling so 10k clients do not share a clock.
This post is that delay. Families and the catalog live on the Algorithms Roadmap. Token Bucket and Leaky Bucket shape a healthy send rate. This formula shapes the failure wait. It is not a circuit-breaker product and not two-phase commit.
The job is a retry delay, not a send rate
A rate limiter answers “may I send this request while the path is healthy?” Token bucket allows a burst, then enforces an average. Leaky bucket drains at a constant rate. Both meter traffic you already intend to send.
Backoff answers “the last call failed — how long until I try the same work again?” The input is an attempt index. The output is a sleep. Substituting one for the other is how a retry storm looks like a traffic spike: the limiter never saw a failure, and the retry loop never spread the wake-ups.
Do not ship a rate limiter and call the outage handled. You often want both — limiter on the healthy path, backoff on the miss. This page only computes the wait.
Note: Stop retrying. A max-attempt (or a deadline) is part of the policy. An uncapped loop is not patience; it is a background flood.
Double, then cap
After attempt attempt — 0 is the first wait after the failed call — the ceiling is:
min(cap, base × 2^attempt)
base is the first wait. Each further miss doubles. cap is the longest you will ever sleep. Without it, a 100ms base at attempt 20 is about 29 hours.
Without jitter, every client that shares base, cap, and attempt wakes on the same millisecond. Exponential delay only spaces the stampedes. It does not unsynchronize them.
Do not use Math.pow for the doubling. Integer multiply that saturates at cap is the lab. 1L << attempt looks clever until attempt is large: Java masks the shift count, and the wait wraps into nonsense.
A walk: 100ms base, attempts 0..5
Same outage. base = 100ms, cap = 2000ms. Ten thousand clients failed at t=0. First the ceiling with no jitter.
attempt ceiling = min(2000, 100 × 2^attempt) all 10k wake at
0 100 t = 100ms
1 200 t = 300ms (100 + 200)
2 400 t = 700ms
3 800 t = 1500ms
4 1600 t = 3100ms
5 min(2000, 3200) = 2000 t = 5100ms
Every client that is still failing shares those instants. The recovered box sees 10k hits at 100ms, then 10k at 300ms, then 10k at 700ms. The wait grew. The herd did not leave.
Full jitter draws uniform in [0, ceiling]. Equal jitter draws in [ceiling/2, ceiling].
attempt ceiling full jitter equal jitter
0 100 [0, 100] [50, 100]
1 200 [0, 200] [100, 200]
2 400 [0, 400] [200, 400]
3 800 [0, 800] [400, 800]
4 1600 [0, 1600] [800, 1600]
5 2000 [0, 2000] [1000, 2000]
At attempt 0 with full jitter, the 10k waits smear across a 100ms window instead of a single tick. Later attempts smear across a wider window — that is the exponential doing the spreading, jitter doing the desynchronizing.
A single client still sees a growing wait. The crowd no longer shares a clock.
Full jitter vs equal jitter
Full jitter is the usual default when the failure mode is a herd. Sleep random(0, ceiling). Maximum spread. A draw near 0 retries almost immediately — that is the cost of unsynchronizing.
Equal jitter sleeps ceiling/2 + random(0, ceiling/2). You never retry faster than half the exponential. Use it when a too-eager retry is worse than a slightly tighter cluster (a write you would rather not duplicate in the first 50ms).
Decorrelated jitter (next wait depends on the previous wait, not on 2^attempt) is a different curve. This post does not walk it. Linear backoff (base × attempt) is also a different curve — slower ramp, still a herd without jitter.
Note: If the server sent Retry-After, honor that wait. Do not invent a shorter full-jitter draw and call it polite.
Java: the delay, not a resilience library
No Resilience4j tutorial. Compute the ceiling, draw jitter with ThreadLocalRandom, sleep. The caller owns max attempts, which exceptions are retryable, and whether the work is idempotent.
static long ceilingMs(int attempt, long baseMs, long capMs) {
if (attempt < 0 || baseMs <= 0 || capMs <= 0) {
throw new IllegalArgumentException("attempt >= 0; base and cap > 0");
}
long delay = Math.min(capMs, baseMs);
for (int i = 0; i < attempt; i++) {
if (delay >= capMs) {
return capMs;
}
delay = Math.min(capMs, delay * 2);
}
return delay;
}
static long fullJitterMs(long ceilingMs) {
if (ceilingMs <= 0) {
return 0;
}
return ThreadLocalRandom.current().nextLong(ceilingMs + 1);
}
static long equalJitterMs(long ceilingMs) {
long half = ceilingMs / 2;
return half + ThreadLocalRandom.current().nextLong(half + 1);
}
nextLong(n) is 0 .. n-1. nextLong(ceilingMs + 1) includes the cap. nextLong(0) throws — that is the zero-ceiling guard.
A retry loop is the delay plus a stop:
for (int attempt = 0; attempt <= 5; attempt++) {
try {
return call();
} catch (RetryableException e) {
if (attempt == 5) {
throw e;
}
long wait = fullJitterMs(ceilingMs(attempt, 100, 2_000));
try {
Thread.sleep(wait);
} catch (InterruptedException ie) {
Thread.currentThread().interrupt();
throw new IllegalStateException(ie);
}
}
}
Note: Saturate at cap before delay * 2 when delay > Long.MAX_VALUE / 2. Realistic caps (seconds to a few minutes) never get there — the >= cap return fires first. Restore the interrupt flag; swallowing InterruptedException hides shutdown.
When not to use this wait
Skip this formula when the job is not “sleep after a failure, then try the same work again.”
- You needed a rate limiter. Healthy-path bursts and averages are Token Bucket and Leaky Bucket. Backoff does not meter success.
- The call is not idempotent. Retrying a POST that already charged a card is a duplicate, not a recovery. An idempotency key is a protocol; it is not this sleep.
- The server named the wait.
Retry-Afterand a 429 with a delay beat a client-side draw. Honor the longer of the two, or the server’s number alone. - You wanted a circuit breaker. Opening a shared “stop calling” is a health signal across requests. This page is a per-attempt delay. Do not fake a breaker by sleeping in one thread.
- You wanted two-phase commit. Prepare-then-commit is a different Wave 7 job. A retry loop is not a transaction protocol.
- Linear was the spec.
base × attemptis not this curve. Implement what the SLA named.
Exponential without jitter still herds — the crowd is just slower between stampedes. Circuit breakers as a product, token/leaky labs, and 2PC are other pages. Do not open them here.
Cheat sheet
Job: sleep after a failure so retries do not stampede
Ceiling: min(cap, base × 2^attempt) (attempt 0 = first retry wait)
Jitter: full = random[0, ceiling]
equal = ceiling/2 + random[0, ceiling/2]
Herd: same attempt + no jitter → same wake-up (10k at t=1s)
Not this: token bucket, leaky bucket, circuit breaker, 2PC
Stop: max attempts or a deadline — the cap is the sleep, not the loop
JDK: ThreadLocalRandom; no Resilience4j tutorial on this page
Cost: O(1) (tiny loop to the cap) to compute; the wait is the bill
Do:
- Cap the ceiling. Double with saturating multiply, not
Math.pow. - Add full jitter unless you have a reason to keep a floor (equal jitter).
- Bound the loop. Honor
Retry-Afterwhen the server sent it. - Keep this a retry policy. Point rate limits at token/leaky.
Don’t:
- Retry 10k clients at the same instant and call exponential “done.”
- Use backoff as a rate limiter, or a limiter as an outage plan.
- Retry forever, or retry a non-idempotent write without a key.
- Hand-roll Resilience4j. Compute the delay; stop at the library boundary.
Wrap-up
Exponential backoff is a wait: min(cap, base × 2^attempt), then a random draw inside that ceiling. Without jitter, ten thousand clients who failed together wake together — the recovered box dies of retries. Full jitter smears the herd; equal jitter keeps a floor. Token bucket and leaky bucket meter the healthy path. This formula does not. Cap the sleep, cap the loop, and leave 2PC and circuit breakers to the jobs they own.
The outage already happened. The procedure is this delay. When the next job is “every participant prepares, then every participant commits,” that is two-phase commit — and it blocks.