Building a Two-Tier Local and Redis Cache¶
A cache that lives only in process memory is fast and lonely: every worker warms its own copy, a deploy empties all of them, and eight workers hold eight versions of the truth. A cache that lives only in Redis is shared and durable, and puts a network round trip on every read. Most production services end up with both — a small local tier in front of Redis — and the design question is how the tiers interact. Measured on a loopback Redis 7.4 with a 20 ms origin: a local hit cost 0.2 µs at the median, a Redis hit 66 µs, and an origin call 20.5 ms. Replaying 20,000 skewed reads over a few hundred keys, the first worker made 148 origin calls and served the other 19,852 locally; a second worker starting cold made zero origin calls in its first 5,000 reads — 74 came from Redis, the rest from its own local tier once warmed.
Prerequisites¶
- Python 3.11+,
pip install redis(theredis.asyncioclient), and a Redis server. - Redis client basics, from caching with redis.asyncio clients.
- Stampede protection, from preventing cache stampedes in asyncio.
1. Read local, then Redis, then origin¶
The read path checks each tier in order of cost and fills the cheaper tiers on the way back:
import json
import time
import redis.asyncio as aioredis
class TwoTierCache:
def __init__(self, r: aioredis.Redis, *, local_ttl: float = 5.0,
remote_ttl: int = 60, maxsize: int = 10_000) -> None:
self.r = r
self.local: dict[str, tuple[float, object]] = {}
self.local_ttl, self.remote_ttl, self.maxsize = local_ttl, remote_ttl, maxsize
self.stats = {"local": 0, "remote": 0, "origin": 0}
async def get(self, key: str, load):
now = time.monotonic()
hit = self.local.get(key)
if hit is not None and hit[0] > now:
self.stats["local"] += 1
return hit[1]
raw = await self.r.get(key)
if raw is not None:
value = json.loads(raw)
self.stats["remote"] += 1
else:
value = await load(key)
self.stats["origin"] += 1
await self.r.set(key, json.dumps(value), ex=self.remote_ttl)
if len(self.local) >= self.maxsize:
self.local.pop(next(iter(self.local))) # evict oldest insertion
self.local[key] = (now + self.local_ttl, value)
return value
The tiers serve different traffic. The local tier absorbs hot keys within one worker; Redis absorbs warm keys across workers and survives restarts; the origin sees each key roughly once per Redis TTL across the whole fleet.
Verify: count hits per tier over a realistic replay; most reads should be local, and origin calls should be close to the number of distinct keys per Redis TTL.
2. Keep the local TTL much shorter than Redis's¶
The local tier is the one that goes stale: an update written to Redis is invisible to a worker that still holds the old value locally. The simplest consistency rule is to make the local TTL short — seconds — and the Redis TTL long. Staleness is then bounded by the local TTL, no matter what else goes wrong:
| Tier | Typical TTL | Bounded by |
|---|---|---|
| local, per worker | 1–10 s | how stale a read may be |
| Redis, shared | minutes to hours | how often the origin may be called |
With a 5 s local TTL, a worker re-reads a hot key from Redis at most once every 5 s — cheap at 66 µs — and any update is visible everywhere within 5 s. If some data cannot be 5 s stale (balances, permissions), it should skip the local tier entirely rather than get a special shorter TTL; per-key exceptions are where cache bugs hide.
Verify: update a value in Redis; every worker serves the new value within one local TTL.
3. Invalidate the local tier on writes¶
When a short TTL is not enough, push invalidations. Publish the key on a Redis channel after each write, and have every worker drop its local copy:
CHANNEL = "cache:invalidate"
async def write_through(cache: TwoTierCache, key: str, value) -> None:
await cache.r.set(key, json.dumps(value), ex=cache.remote_ttl)
cache.local.pop(key, None) # this worker, immediately
await cache.r.publish(CHANNEL, key) # every other worker
async def invalidation_listener(cache: TwoTierCache) -> None:
async with cache.r.pubsub() as ps:
await ps.subscribe(CHANNEL)
async for msg in ps.listen():
if msg["type"] == "message":
cache.local.pop(msg["data"].decode(), None)
Pub/sub is fire-and-forget: a worker whose subscription is reconnecting misses the message and keeps its stale copy until the local TTL expires. That is why the short TTL stays even with invalidation — it is the backstop. The failure behaviour across workers, measured, is in invalidating caches across async workers.
Verify: after a write, other workers drop the key within milliseconds; with the listener stopped, they still converge within the local TTL.
4. Deduplicate misses in both tiers¶
Two kinds of stampede apply. Within a worker, many concurrent requests for a key that is in neither tier would all call the origin; across workers, every worker can miss Redis at once. Put a per-key in-flight task in front of the slow path, so concurrent local misses share one Redis read and, if needed, one origin call:
class DedupedTwoTier(TwoTierCache):
def __init__(self, *a, **kw) -> None:
super().__init__(*a, **kw)
self._inflight: dict[str, asyncio.Task] = {}
async def get(self, key: str, load):
hit = self.local.get(key)
if hit is not None and hit[0] > time.monotonic():
self.stats["local"] += 1
return hit[1]
task = self._inflight.get(key)
if task is None:
task = asyncio.ensure_future(super().get(key, load))
self._inflight[key] = task
task.add_done_callback(lambda _: self._inflight.pop(key, None))
return await asyncio.shield(task)
This handles the in-process half. The cross-worker half — every worker missing Redis at the same moment after a popular key expires — needs either a Redis lock around the origin call or stale-while-revalidate, both covered in the stampede guide linked above. For most keys, the local tier plus per-worker deduplication already reduces origin calls to one per worker per expiry.
Verify: fire 100 concurrent reads for an uncached key in one worker; Redis sees one GET, the origin one call.
5. Serialise values deliberately¶
Every Redis hit pays for deserialisation, and every miss for serialisation. JSON is portable and slow for large objects; pickle is faster for Python objects and unsafe if anything untrusted can write to Redis. Measure with your real values:
import pickle
PREFIX_JSON, PREFIX_PICKLE = b"j", b"p"
def dumps(value) -> bytes:
try:
return PREFIX_JSON + json.dumps(value, separators=(",", ":")).encode()
except TypeError:
return PREFIX_PICKLE + pickle.dumps(value, protocol=pickle.HIGHEST_PROTOCOL)
def loads(raw: bytes):
return json.loads(raw[1:]) if raw[:1] == PREFIX_JSON else pickle.loads(raw[1:])
A one-byte prefix lets the format change without flushing Redis: old entries keep decoding while new writes use the new format. Version the key namespace (v2:user:42) when the shape of the value changes, so workers on old and new code during a deploy do not read each other's incompatible objects.
Verify: during a rolling deploy that changes a cached type, no worker raises a decode error.
Verification¶
The two-tier cache is correct when:
- Hit distribution matches expectations: most reads local, origin calls near distinct keys per Redis TTL.
- Staleness is bounded by the local TTL even when invalidation messages are lost.
- Concurrent misses are deduplicated within each worker.
- Format and shape changes are handled by prefixes and versioned key namespaces.
Diagnostic Hook: export per-tier hit counters and latency histograms, plus local tier size and invalidations received. The ratio that matters most is origin calls per minute against distinct keys per minute — if it rises above roughly one call per key per Redis TTL, either the Redis TTL is too short or keys are being evicted by Redis's memory policy, which INFO stats reports as evicted_keys.
Pitfalls & edge cases¶
- Local TTL as long as Redis's. Updates then take the full TTL to appear; keep the local tier short-lived.
- Unbounded local tier. A dict keyed by user input is a memory leak; cap it and evict.
- Caching mutable objects locally. Callers that mutate the returned dict corrupt everyone's copy; return copies or immutable values.
- pickle with untrusted writers. Anyone who can write to Redis can run code in your workers.
Frequently Asked Questions¶
Why use a local cache in front of Redis?
A local hit cost 0.2 µs against 66 µs for a Redis hit in testing. For hot keys read many times per second, the local tier removes almost all Redis round trips while Redis still shares warm data across workers and survives deploys.
How do I keep a local cache consistent with Redis?
Give the local tier a short TTL, seconds, so staleness is bounded, and optionally publish invalidation messages on writes so workers drop their copies immediately. The TTL remains the backstop for missed messages.
How long should the local and Redis TTLs be?
Set the local TTL from how stale a read may be, typically 1 to 10 seconds, and the Redis TTL from how often the origin may be called, typically minutes or longer.
How do I stop many workers calling the origin at once?
Deduplicate concurrent misses within each worker with an in-flight task per key, and for very hot keys add a Redis lock or stale-while-revalidate so only one worker rebuilds an expired entry.
Related¶
- Async Caching & Deduplication — up to the topic overview.
- Bounding in-memory async caches by size — sizing the local tier.
- Concurrent Execution & Worker Patterns — the section overview.