Skip to content

Refreshing Cache Entries Ahead of Expiry

A cache with a plain TTL makes someone pay for every expiry: the first request after an entry expires waits for the origin, and so does every request that arrives while that fetch is in flight. For hot keys this happens once per TTL per key, forever. Refresh-ahead moves the fetch earlier: when a request hits an entry in the last part of its lifetime, the cache starts a background refresh and answers immediately from the current value, so popular entries are replaced before anyone sees them expire. Measured on Python 3.14 with 50 hot keys, a 1-second TTL, an origin taking 50 ms and 2,000 requests per second for 8 s: with a plain TTL, 1,161 requests waited on the origin and p99 latency was 50.7 ms. Refreshing in the background once an entry passed 80% of its TTL cut waiting requests to 202 — mostly the first request for each key — and p99 to 20.0 ms, at the cost of 457 origin calls instead of 382, about 20% more. Across 1,000 keys with a long tail of rarely requested ones, the effect was smaller — 3,532 waits fell to 2,629 — because refresh-ahead only helps keys that are requested during their refresh window. This guide implements it and shows where it pays.

Prerequisites

1. Count the requests that wait on expiry

The cost of expiry is not the origin call itself but the requests that wait for it. Count requests that had to await a fetch, separately from hits:

class TTLCache:
    async def get(self, key):
        entry = self.data.get(key)
        if entry and self._now() - entry.fetched_at < self.ttl:
            return entry.value
        self.waited += 1                                  # this request pays origin latency
        return await asyncio.shield(self._start_load(key))

Measured with 50 hot keys and a 1-second TTL: 1,161 of 16,000 requests waited on the origin — the first request after each expiry plus everything that arrived during the 50 ms fetch — and p99 latency was 50.7 ms, the origin's latency, because more than 1% of requests waited. Coalescing already kept it to one origin call per expiry, 382 calls in all; what remained was the waiting.

Verify: you know what fraction of requests wait on the origin, and how much of it is expiry rather than first access.

2. Start a background refresh in the last part of an entry's life

On a hit, check the entry's age; past a threshold, start a refresh without awaiting it, and return the current value:

class RefreshAheadCache(TTLCache):
    def __init__(self, origin, ttl: float = 60.0, refresh_at: float = 0.8):
        super().__init__(origin, ttl)
        self.refresh_at = refresh_at

    async def get(self, key):
        entry = self.data.get(key)
        now = self._now()
        if entry and now - entry.fetched_at < self.ttl:
            if now - entry.fetched_at > self.refresh_at * self.ttl:
                self._start_load(key)                     # coalesced; not awaited
            return entry.value
        self.waited += 1
        return await asyncio.shield(self._start_load(key))

_start_load is the same coalescing loader the cache uses for misses, so a hot key gets one background refresh, not one per hit, and the refreshed value replaces the entry when it arrives. Measured: requests that waited on the origin fell from 1,161 to 202, and p99 latency from 50.7 ms to 20.0 ms. The remaining waits were each key's first request and keys that were not requested during their refresh window. Keep a reference to background loads — the in-flight map does that here — and log their failures, since no caller is waiting to see them, as discussed in handling exceptions in fire-and-forget tasks.

Verify: for hot keys, requests after warm-up almost never wait on the origin.

16,000 requests over 8 s, 1 s TTL, 50 ms origin A grid of 4 rows by 5 columns. 16,000 requests over 8 s, 1 s TTL, 50 ms origin keys cache requests that waited p99 origin calls 50 hot plain TTL 1,161 50.7 ms 382 50 hot refresh-ahead at 80% 202 20.0 ms 457 1,000, Zipf tail plain TTL 3,532 51.1 ms 2,879 1,000, Zipf tail refresh-ahead at 80% 2,629 51.1 ms 3,161 Python 3.14; refreshes coalesced per key.

3. Pay for refreshes only where traffic justifies them

Refresh-ahead refreshes an entry because it was requested near its expiry — not because it will be requested again. For keys requested just once in their refresh window, the refresh is wasted:

REFRESH_AT = 0.8     # refresh during the last 20% of the TTL

Measured: origin calls rose from 382 to 457 for 50 hot keys, about 20% more, because some refreshed entries expired again unused. Across 1,000 keys with a long tail, refresh-ahead cut waiting requests only from 3,532 to 2,629 and did not change p99 — rarely requested keys were usually cold or expired when asked for, and no refresh window could help them. A later threshold (0.9) spends fewer extra calls and protects fewer requests; an earlier one (0.5) does the reverse. Tune it on your key distribution, and consider refreshing only keys with a minimum request rate, tracked with a small counter per entry.

Verify: extra origin calls from refreshes are measured, and the threshold is chosen against the reduction in waiting requests.

A hot key under refresh-ahead A sequence of 6 messages between 4 participants. A hot key under refresh-ahead request cache background load origin get: age 0.5 s, hit get: age 0.85 s, hit start refresh (not awaited) fetch, 50 ms replace entry at age 0.9 s get: fresh again, hit The entry never expires while it keeps being requested.

4. Spread refreshes so they do not synchronise

Entries loaded together expire together, and with refresh-ahead they also refresh together — a burst of origin calls every TTL. Add jitter to the refresh threshold per entry, so a set of keys warmed at start-up drifts apart over time:

import random

def refresh_deadline(fetched_at: float, ttl: float, refresh_at: float = 0.8, jitter: float = 0.1) -> float:
    return fetched_at + ttl * (refresh_at - random.uniform(0, jitter))

Store the deadline with the entry when it is loaded and compare against it on each hit. Bound concurrent background refreshes too, with a semaphore around the origin call: when the origin slows down, refreshes otherwise pile up behind it at exactly the moment it can least afford extra load. Cache warming at start-up, which creates the synchronised batch in the first place, is covered in warming caches on startup without blocking readiness.

Verify: origin call rate from refreshes is smooth over time rather than spiking once per TTL.

5. Decide between refresh-ahead and stale-while-revalidate

Refresh-ahead never serves anything older than the TTL, but it cannot protect a key that is not requested in its refresh window, and it does nothing when the origin is down. Stale-while-revalidate serves the expired value while refreshing, which removed nearly all waiting in the same test — 148 waiting requests and a p99 of 0.0 ms — and, during a two-second origin outage, answered every request from stale data while the plain-TTL and refresh-ahead caches returned 3,509 and 2,611 errors. The difference is what you promise about freshness:

# refresh-ahead: values are never older than ttl; some requests still wait
cache = RefreshAheadCache(origin, ttl=60, refresh_at=0.8)

# stale-while-revalidate: values may be up to ttl + grace old; almost nobody waits
cache = SWRCache(origin, ttl=60, stale_for=300)

Use refresh-ahead where data must never be older than its TTL — prices, permissions, anything with a correctness deadline. Use stale-while-revalidate, described in serving stale while revalidating, where a slightly old answer is better than a slow one or an error.

Verify: each cache's freshness guarantee is written down and matches the technique it uses.

Which expiry strategy fits this data? A decision on What may a reader see with 4 outcomes. Which expiry strategy fits this data? What may a reader see? never older than TTL, hot keys refresh-ahead waits 1,161 to 202 slightly stale is fine stale-while-revalidate waits 1,161 to 148 long tail of rare keys either helps little they miss anyway many keys loaded together jitter + bounded refreshes no synchronised bursts Refresh-ahead trades extra origin calls for fewer waiting requests.

Verification

Refresh-ahead is working when:

  • Waiting requests are counted, and drop for hot keys after the change.
  • Background refreshes are coalesced per key, referenced, and their errors logged.
  • Extra origin calls are measured and justified by the reduction in waiting.
  • Refresh deadlines are jittered and concurrent refreshes bounded.

Diagnostic Hook: export three counters — hits, waits, and background refreshes — per cache. Refreshes far above waits saved means the threshold is too early or the keys too cold; waits that stay high for keys you know are hot means refreshes are failing silently in the background.

Pitfalls & edge cases

  • Expecting help for cold keys. Measured: 3,532 waits became 2,629, not zero, with a long tail.
  • Uncoalesced background refreshes. Every hit in the window would call the origin.
  • Unlogged refresh failures. The cache keeps serving until the entry expires, then everyone waits.
  • Synchronised expiry. Entries loaded together refresh together without jitter.

Frequently Asked Questions

What is refresh-ahead caching?

Refreshing an entry in the background when it is requested near the end of its TTL, while returning the current value. For 50 hot keys it cut requests that waited on the origin from 1,161 to 202 in testing.

Does refresh-ahead increase load on the origin?

Somewhat: origin calls rose from 382 to 457, about 20%, because some refreshed entries expired again unused.

When should refresh-ahead trigger?

In the last part of the TTL; 80% was used here. Later triggers waste fewer calls and protect fewer requests. Add jitter so entries loaded together do not refresh together.

Is refresh-ahead better than stale-while-revalidate?

It never serves data older than the TTL, but stale-while-revalidate removed more waiting (148 requests) and survived an origin outage that caused 2,611 errors with refresh-ahead.