Refreshing Cache Entries Ahead of Expiry¶
A cache with a plain TTL makes someone pay for every expiry: the first request after an entry expires waits for the origin, and so does every request that arrives while that fetch is in flight. For hot keys this happens once per TTL per key, forever. Refresh-ahead moves the fetch earlier: when a request hits an entry in the last part of its lifetime, the cache starts a background refresh and answers immediately from the current value, so popular entries are replaced before anyone sees them expire. Measured on Python 3.14 with 50 hot keys, a 1-second TTL, an origin taking 50 ms and 2,000 requests per second for 8 s: with a plain TTL, 1,161 requests waited on the origin and p99 latency was 50.7 ms. Refreshing in the background once an entry passed 80% of its TTL cut waiting requests to 202 — mostly the first request for each key — and p99 to 20.0 ms, at the cost of 457 origin calls instead of 382, about 20% more. Across 1,000 keys with a long tail of rarely requested ones, the effect was smaller — 3,532 waits fell to 2,629 — because refresh-ahead only helps keys that are requested during their refresh window. This guide implements it and shows where it pays.
Prerequisites¶
- Python 3.11+; an async cache with request coalescing.
- Coalescing concurrent misses, from preventing cache stampedes in asyncio.
- The topic overview, Async Caching & Deduplication.
1. Count the requests that wait on expiry¶
The cost of expiry is not the origin call itself but the requests that wait for it. Count requests that had to await a fetch, separately from hits:
class TTLCache:
async def get(self, key):
entry = self.data.get(key)
if entry and self._now() - entry.fetched_at < self.ttl:
return entry.value
self.waited += 1 # this request pays origin latency
return await asyncio.shield(self._start_load(key))
Measured with 50 hot keys and a 1-second TTL: 1,161 of 16,000 requests waited on the origin — the first request after each expiry plus everything that arrived during the 50 ms fetch — and p99 latency was 50.7 ms, the origin's latency, because more than 1% of requests waited. Coalescing already kept it to one origin call per expiry, 382 calls in all; what remained was the waiting.
Verify: you know what fraction of requests wait on the origin, and how much of it is expiry rather than first access.
2. Start a background refresh in the last part of an entry's life¶
On a hit, check the entry's age; past a threshold, start a refresh without awaiting it, and return the current value:
class RefreshAheadCache(TTLCache):
def __init__(self, origin, ttl: float = 60.0, refresh_at: float = 0.8):
super().__init__(origin, ttl)
self.refresh_at = refresh_at
async def get(self, key):
entry = self.data.get(key)
now = self._now()
if entry and now - entry.fetched_at < self.ttl:
if now - entry.fetched_at > self.refresh_at * self.ttl:
self._start_load(key) # coalesced; not awaited
return entry.value
self.waited += 1
return await asyncio.shield(self._start_load(key))
_start_load is the same coalescing loader the cache uses for misses, so a hot key gets one background refresh, not one per hit, and the refreshed value replaces the entry when it arrives. Measured: requests that waited on the origin fell from 1,161 to 202, and p99 latency from 50.7 ms to 20.0 ms. The remaining waits were each key's first request and keys that were not requested during their refresh window. Keep a reference to background loads — the in-flight map does that here — and log their failures, since no caller is waiting to see them, as discussed in handling exceptions in fire-and-forget tasks.
Verify: for hot keys, requests after warm-up almost never wait on the origin.
3. Pay for refreshes only where traffic justifies them¶
Refresh-ahead refreshes an entry because it was requested near its expiry — not because it will be requested again. For keys requested just once in their refresh window, the refresh is wasted:
REFRESH_AT = 0.8 # refresh during the last 20% of the TTL
Measured: origin calls rose from 382 to 457 for 50 hot keys, about 20% more, because some refreshed entries expired again unused. Across 1,000 keys with a long tail, refresh-ahead cut waiting requests only from 3,532 to 2,629 and did not change p99 — rarely requested keys were usually cold or expired when asked for, and no refresh window could help them. A later threshold (0.9) spends fewer extra calls and protects fewer requests; an earlier one (0.5) does the reverse. Tune it on your key distribution, and consider refreshing only keys with a minimum request rate, tracked with a small counter per entry.
Verify: extra origin calls from refreshes are measured, and the threshold is chosen against the reduction in waiting requests.
4. Spread refreshes so they do not synchronise¶
Entries loaded together expire together, and with refresh-ahead they also refresh together — a burst of origin calls every TTL. Add jitter to the refresh threshold per entry, so a set of keys warmed at start-up drifts apart over time:
import random
def refresh_deadline(fetched_at: float, ttl: float, refresh_at: float = 0.8, jitter: float = 0.1) -> float:
return fetched_at + ttl * (refresh_at - random.uniform(0, jitter))
Store the deadline with the entry when it is loaded and compare against it on each hit. Bound concurrent background refreshes too, with a semaphore around the origin call: when the origin slows down, refreshes otherwise pile up behind it at exactly the moment it can least afford extra load. Cache warming at start-up, which creates the synchronised batch in the first place, is covered in warming caches on startup without blocking readiness.
Verify: origin call rate from refreshes is smooth over time rather than spiking once per TTL.
5. Decide between refresh-ahead and stale-while-revalidate¶
Refresh-ahead never serves anything older than the TTL, but it cannot protect a key that is not requested in its refresh window, and it does nothing when the origin is down. Stale-while-revalidate serves the expired value while refreshing, which removed nearly all waiting in the same test — 148 waiting requests and a p99 of 0.0 ms — and, during a two-second origin outage, answered every request from stale data while the plain-TTL and refresh-ahead caches returned 3,509 and 2,611 errors. The difference is what you promise about freshness:
# refresh-ahead: values are never older than ttl; some requests still wait
cache = RefreshAheadCache(origin, ttl=60, refresh_at=0.8)
# stale-while-revalidate: values may be up to ttl + grace old; almost nobody waits
cache = SWRCache(origin, ttl=60, stale_for=300)
Use refresh-ahead where data must never be older than its TTL — prices, permissions, anything with a correctness deadline. Use stale-while-revalidate, described in serving stale while revalidating, where a slightly old answer is better than a slow one or an error.
Verify: each cache's freshness guarantee is written down and matches the technique it uses.
Verification¶
Refresh-ahead is working when:
- Waiting requests are counted, and drop for hot keys after the change.
- Background refreshes are coalesced per key, referenced, and their errors logged.
- Extra origin calls are measured and justified by the reduction in waiting.
- Refresh deadlines are jittered and concurrent refreshes bounded.
Diagnostic Hook: export three counters — hits, waits, and background refreshes — per cache. Refreshes far above waits saved means the threshold is too early or the keys too cold; waits that stay high for keys you know are hot means refreshes are failing silently in the background.
Pitfalls & edge cases¶
- Expecting help for cold keys. Measured: 3,532 waits became 2,629, not zero, with a long tail.
- Uncoalesced background refreshes. Every hit in the window would call the origin.
- Unlogged refresh failures. The cache keeps serving until the entry expires, then everyone waits.
- Synchronised expiry. Entries loaded together refresh together without jitter.
Frequently Asked Questions¶
What is refresh-ahead caching?
Refreshing an entry in the background when it is requested near the end of its TTL, while returning the current value. For 50 hot keys it cut requests that waited on the origin from 1,161 to 202 in testing.
Does refresh-ahead increase load on the origin?
Somewhat: origin calls rose from 382 to 457, about 20%, because some refreshed entries expired again unused.
When should refresh-ahead trigger?
In the last part of the TTL; 80% was used here. Later triggers waste fewer calls and protect fewer requests. Add jitter so entries loaded together do not refresh together.
Is refresh-ahead better than stale-while-revalidate?
It never serves data older than the TTL, but stale-while-revalidate removed more waiting (148 requests) and survived an origin outage that caused 2,611 errors with refresh-ahead.
Related¶
- Async Caching & Deduplication — up to the topic overview.
- Serving stale while revalidating — trading freshness for availability.
- Concurrent Execution & Worker Patterns — the section overview.