Skip to content

Serving Stale While Revalidating

Stale-while-revalidate lets a cache answer with a value that has just expired, while one background fetch replaces it. Readers almost never wait for the origin, and when the origin is down, they keep getting the last good answer instead of an error. The price is a bound on staleness that is looser than the TTL, which has to be chosen deliberately. Measured on Python 3.14 with 50 hot keys, a 1-second TTL, an origin taking 50 ms and 2,000 requests per second for 8 s: a plain TTL cache made 1,161 requests wait on the origin with p99 50.7 ms; serving stale entries for up to 5 s after expiry made 148 wait — essentially each key's first request — with p99 0.0 ms, the same 382 origin calls, and 1,013 responses served from just-expired entries. During a 2-second origin outage, the plain cache returned 3,509 errors and stale-while-revalidate none. With a grace period of only 1 s and a 5-second outage, the cache returned 7,615 errors once entries aged past the grace; adding a stale-if-error allowance answered all 7,623 of those requests from data up to 6.05 s old, with 0 errors. This guide builds that cache.

Prerequisites

1. Define three ages: fresh, stale, and too old

Each entry has an age since it was fetched. Below the TTL it is fresh and served as is; for a grace period after that it is stale, served immediately while a refresh runs; beyond the grace, the request waits for the origin like a normal miss:

class SWRCache:
    def __init__(self, origin, ttl: float = 60.0, stale_for: float = 300.0, stale_if_error: float = 3600.0):
        self.origin = origin
        self.ttl, self.stale_for, self.stale_if_error = ttl, stale_for, stale_if_error
        self.data: dict = {}                 # key -> (value, fetched_at)
        self.inflight: dict = {}

    async def get(self, key):
        entry = self.data.get(key)
        now = self._now()
        if entry:
            age = now - entry[1]
            if age < self.ttl:
                return entry[0]                                  # fresh
            if age < self.ttl + self.stale_for:
                self._start_load(key)                            # refresh in the background
                return entry[0]                                  # answer with the stale value now
        try:
            return await asyncio.shield(self._start_load(key))   # missing or too old: wait
        except (ConnectionError, TimeoutError):
            if entry and now - entry[1] < self.ttl + self.stale_if_error:
                return entry[0]                                  # origin failed: stale is better than nothing
            raise

_start_load coalesces per key, so however many requests hit a stale entry at once, the origin sees one refresh. The three numbers are the cache's contract: values are fresh for ttl, may be up to ttl + stale_for old on any normal request, and up to ttl + stale_if_error old while the origin is failing.

Verify: each cached data type has documented values for the three ages.

2. Measure what readers stop waiting for

Under normal operation, stale-while-revalidate removes the wait at expiry entirely for keys that are requested within their grace period:

cache = SWRCache(origin, ttl=1.0, stale_for=5.0)

Measured with 50 hot keys: 148 of 16,000 requests waited on the origin — one first request per key plus a few — against 1,161 for a plain TTL and 202 for refresh-ahead; p99 latency was 0.0 ms against 50.7 ms; and origin calls were 382, the same as the plain cache, because each expiry still caused exactly one fetch. 1,013 requests were answered from entries that had expired, by at most the fetch time — a refresh took about 50 ms, so in steady state the extra staleness readers saw was about one origin round trip. Compared with refresh-ahead, there are no speculative refreshes: a key that nobody asks for again is never refetched.

Verify: after warm-up, the share of requests that wait on the origin is close to the share of first-time keys.

50 hot keys, 1 s TTL, 16,000 requests A grid of 3 rows by 5 columns. 50 hot keys, 1 s TTL, 16,000 requests cache requests that waited p99 origin calls errors in a 2 s outage plain TTL 1,161 50.7 ms 382 3,509 refresh-ahead at 80% 202 20.0 ms 457 2,611 stale-while-revalidate, 5 s grace 148 0.0 ms 382 0 Origin latency 50 ms; outage from 3 s to 5 s of an 8 s run.

3. Keep serving through origin outages, within a limit

When the origin fails, a stale entry within its grace still answers immediately, and its background refresh fails quietly. Measured during a 2-second outage with a 5-second grace: 4,263 requests were answered from stale entries and none failed, while the plain cache returned 3,509 errors and refresh-ahead 2,611. If the outage outlasts the grace, entries age past it and requests fall through to the failing origin:

cache = SWRCache(origin, ttl=1.0, stale_for=1.0, stale_if_error=0.0)    # 7,615 errors
cache = SWRCache(origin, ttl=1.0, stale_for=1.0, stale_if_error=60.0)   # 0 errors

Measured with a 1-second grace and a 5-second outage: without a stale-if-error allowance, 7,615 requests failed; with 60 seconds of allowance, the same 7,623 requests were answered from entries up to 6.05 s old, and none failed. Those requests still waited for the failed origin attempt before falling back — coalesced per key, so one attempt at a time — which a circuit breaker in front of the origin can short-circuit, as in Circuit Breakers & Bulkheads. Set stale_if_error from how old an answer may be before it is worse than an error: minutes for a product catalogue, seconds or zero for account balances.

Verify: during an injected origin outage, errors appear only after entries exceed ttl + stale_if_error.

A request during an origin outage A sequence of 5 messages between 3 participants. A request during an origin outage request cache origin get: age 2.5 s (> ttl + grace) fetch (coalesced) ConnectionError age < ttl + stale_if_error stale value, 2.5 s old With stale_if_error = 0, the same request would have failed.

4. Tell callers how old the answer is

Stale data is fine only if code that cares can tell. Return the age with the value, or expose it through response headers, so callers and clients can decide:

@dataclass
class Cached:
    value: object
    age: float
    stale: bool

async def get_with_age(self, key) -> Cached:
    value = await self.get(key)
    age = self._now() - self.data[key][1]
    return Cached(value, age, age >= self.ttl)

# in an HTTP handler
resp.headers["Age"] = str(int(cached.age))
resp.headers["Cache-Control"] = "max-age=60, stale-while-revalidate=300, stale-if-error=3600"

HTTP defines the same model in Cache-Control's stale-while-revalidate and stale-if-error directives, so a service can apply it internally and declare it to clients and CDNs with the same numbers, as discussed in caching HTTP responses with ETags in async clients. Log when stale-if-error answers are served — they are invisible otherwise, and an origin that has been failing for an hour behind a cache that hides it is a problem you want to hear about before the allowance runs out.

Verify: responses expose their age, and stale-if-error answers are counted and alerted on.

5. Bound background refreshes

Every stale hit can start a refresh, and refreshes run without a waiting caller to apply backpressure. When the origin is slow, many keys can be refreshing at once; when it is failing, every stale hit after a failed refresh starts another. Limit concurrent background refreshes, and remember recent failures briefly so a failing key is not retried on every request:

REFRESH_SLOTS = asyncio.Semaphore(20)
RETRY_AFTER = 1.0                                     # seconds before retrying a failed refresh

async def _load(self, key):
    if self._now() - self.failed_at.get(key, -1e9) < RETRY_AFTER:
        raise ConnectionError("recently failed")      # do not hammer a failing origin
    async with REFRESH_SLOTS:
        try:
            value = await self.origin.get(key)
        except Exception:
            self.failed_at[key] = self._now()
            raise
    self.data[key] = (value, self._now())
    self.failed_at.pop(key, None)
    return value

The semaphore keeps refresh load proportional to what the origin can take, and the failure memory turns a burst of retries into one attempt per key per second during an outage. Both matter most when they are least visible — in the background, during an incident.

Verify: during an origin outage, refresh attempts per key are bounded, and the number of concurrent refreshes never exceeds the limit.

How stale may this data be? A decision on What is a stale answer worth with 4 outcomes. How stale may this data be? What is a stale answer worth? nothing: must be current no stale serving use refresh-ahead or no cache fine for a few minutes stale_for = minutes waits 1,161 to 148 better than an error stale_if_error = longest acceptable age 7,615 errors to 0 callers must know expose Age / stale flag log stale-if-error The three ages are the cache's contract with its readers.

Verification

Stale-while-revalidate is in place when:

  • Fresh, stale and stale-if-error ages are documented per data type.
  • Stale hits return at once and trigger one coalesced refresh.
  • Origin failures fall back to stale data within the allowance, and are logged.
  • Background refreshes are bounded and back off after failures, and responses expose their age.

Diagnostic Hook: export the age of every served value as a histogram. In steady state it should sit just above the TTL at most; a tail stretching toward ttl + stale_if_error means the origin has been failing and the cache has been quietly covering for it — the moment to look, before the allowance expires and the errors arrive all at once.

Pitfalls & edge cases

  • A grace period shorter than likely outages. Measured: 7,615 errors once entries aged past it.
  • Serving stale data silently. Callers cannot tell; expose the age.
  • Unbounded background refreshes. They pile up exactly when the origin is slow.
  • Retrying a failing refresh on every hit. Remember failures briefly.

Frequently Asked Questions

What is stale-while-revalidate?

Serving an expired cache entry immediately while one background fetch refreshes it. With 50 hot keys, it cut requests that waited on the origin from 1,161 to 148 and p99 from 50.7 ms to 0.0 ms in testing.

How does stale-while-revalidate help during outages?

Stale entries keep answering while refreshes fail: a 2-second origin outage caused 0 errors, against 3,509 with a plain TTL cache.

What is stale-if-error?

A longer allowance for serving stale data when the origin fails. With a 1-second grace and a 5-second outage, it turned 7,615 errors into 7,623 answers up to 6.05 s old.

Does stale-while-revalidate increase origin load?

No: it made the same 382 origin calls as a plain TTL cache, one per expiry per requested key, without speculative refreshes.