Skip to content

Warming Caches on Startup Without Blocking Readiness

A freshly started worker has an empty local cache, and the first minutes of traffic after a deploy pay for it: every request misses, every miss goes to the origin, and a fleet restarting together sends the origin a load spike exactly when it can least absorb one. Warming the cache before taking traffic fixes the latency cliff — and, done naively, creates its own spike. In a simulation warming 1,000 keys against an origin with 20 ms base latency, an unbounded warm-up finished fastest, in 0.42 s, by putting 1,000 concurrent requests on the origin; a concurrency limit of 50 took 0.82 s with at most 50 in flight; a limit of 10 took 2.46 s. This guide picks that limit deliberately, splits the warm-up into a blocking hot set and a background tail, and keeps readiness honest about which one is done.

Prerequisites

1. Decide what is worth warming

Warm the keys that the first minute of traffic will actually ask for, not everything the cache has ever held. The best source is recent access data — the most-requested keys over the last hour, recorded by a running worker:

import collections


class AccessLog:
    """Counts key reads; a running worker periodically saves the top N for the next start."""

    def __init__(self) -> None:
        self.counts: collections.Counter[str] = collections.Counter()

    def record(self, key: str) -> None:
        self.counts[key] += 1

    async def save_hot_set(self, redis, n: int = 1000) -> None:
        hot = [k for k, _ in self.counts.most_common(n)]
        await redis.set("cache:hot_set", "\n".join(hot), ex=86_400)
        self.counts.clear()

Access patterns are usually heavily skewed, so a few hundred keys cover most early reads. Warming the next ten thousand has diminishing value and real cost on the origin. If a shared Redis tier exists, the new worker's local tier can warm from Redis instead of the origin — 66 µs per key instead of 20 ms, as measured in building a two-tier local and Redis cache.

Verify: the saved hot set covers a large share of the first minute's reads after a restart — measure the hit rate during that minute.

2. Warm with a concurrency limit

import asyncio


async def warm(cache, keys: list[str], load, *, limit: int = 50) -> int:
    sem = asyncio.Semaphore(limit)
    warmed = 0

    async def one(key: str) -> None:
        nonlocal warmed
        async with sem:
            try:
                await cache.get(key, load)
                warmed += 1
            except Exception as exc:                 # one bad key must not abort the warm-up
                log.warning("warm %s failed: %r", key, exc)

    async with asyncio.TaskGroup() as tg:
        for key in keys:
            tg.create_task(one(key))
    return warmed

The limit is the origin's budget for warm-up traffic, not a performance knob. Measured on the simulation: 50 in flight warmed 1,000 keys in 0.82 s; removing the limit halved that by sending all 1,000 at once. Multiply by the number of workers restarting together and the unbounded version is a thousand-request-per-worker burst — the cache stampede the warm-up was meant to prevent, self-inflicted. Size the limit from what the origin can absorb on top of live traffic, divided by the number of workers that start simultaneously.

Catching per-key exceptions inside one() matters: in a TaskGroup, one failing child cancels the rest, and a warm-up should degrade to "mostly warm" rather than "not warm at all".

Verify: the origin's concurrent-request metric during a deploy stays at or below limit × workers starting at once.

Warming 1,000 keys at different concurrency limits 3 horizontal bars comparing unbounded, 1,000 in flight with the others. Warming 1,000 keys at different concurrency limits unbounded, 1,000 in flight 0.42 s limit 50 0.82 s limit 10 2.46 s Simulated origin: 20 ms plus 0.4 ms per concurrent request. The fastest warm-up is the one that hits the origin hardest; the limit is a budget, not a tuning knob.

3. Block readiness on the hot set only

Split the warm-up in two. The hot set — the few hundred keys that dominate traffic — is warmed before the worker reports ready. The tail is warmed in the background after the worker starts serving:

async def startup(app) -> None:
    hot = await load_hot_set(app.redis)                     # list of keys
    head, tail = hot[:300], hot[300:]

    try:
        async with asyncio.timeout(10):                      # never block readiness forever
            n = await warm(app.cache, head, app.load, limit=50)
        log.info("warmed %d/%d hot keys before ready", n, len(head))
    except TimeoutError:
        log.warning("hot-set warm-up timed out; starting partially warm")

    app.ready.set()                                          # readiness probe goes green
    app.spawn(warm(app.cache, tail, app.load, limit=10), name="warm:tail")

The timeout is essential. A warm-up that blocks readiness indefinitely turns an origin outage into a fleet that cannot start — exactly when you need fresh workers. A partially warm worker serving traffic is better than a perfectly warm worker that never starts. The tail warms at a lower limit because it competes with live traffic.

Verify: with the origin slowed artificially, workers become ready after the timeout and log the partial warm count.

A worker's first seconds with a split warm-up 3 lanes over time. A worker's first seconds with a split warm-up warm-up hot set, limit 50, max 10 s tail in background, limit 10 readiness not ready ready traffic serving, mostly local hits time → Readiness waits for the keys that matter, never for the whole cache.

4. Stagger warm-ups across the fleet

A rolling deploy restarts workers a few at a time, which spreads the warm-up load naturally. A full-fleet restart — scaling from zero, a node pool replacement, a bad deploy rolled back all at once — does not. Add jitter before the warm-up so simultaneous starts spread out:

import os
import random


async def staggered_warm(app) -> None:
    replicas = int(os.environ.get("EXPECTED_REPLICAS", "1"))
    delay = random.uniform(0, min(5.0, 0.1 * replicas))     # up to 5 s, scaled by fleet size
    await asyncio.sleep(delay)
    await warm(app.cache, app.hot_set, app.load, limit=max(5, 200 // replicas))

Dividing the origin's warm-up budget by the replica count is the fleet-level version of the per-worker limit. If a shared Redis tier exists, the first worker to warm a key from the origin fills Redis for everyone after it, so later workers' warm-ups hit Redis rather than the origin — staggering makes that sharing far more effective. The same jitter logic for periodic jobs is in scheduling cron jobs inside an asyncio service.

Verify: during a full-fleet restart, the origin's request rate rises gradually rather than as one spike.

5. Measure whether warming helps

Warm-up is extra code on the most fragile path a service has, startup. Keep it only if it measurably helps:

async def first_minute_report(cache, started_at: float) -> None:
    await asyncio.sleep(60)
    s = cache.stats
    total = sum(s.values()) or 1
    log.info("first minute after start: local %.0f%%, remote %.0f%%, origin %.0f%%",
             100 * s["local"] / total, 100 * s["remote"] / total, 100 * s["origin"] / total)

Compare the first-minute origin share and p99 latency with and without warming, over several deploys. If the local tier refills within seconds from live traffic anyway — common for services with a small, very hot key set — the warm-up can go. If the first minute shows an origin share several times the steady state and a latency bump, warming earns its complexity.

Verify: you have before-and-after numbers for first-minute p99 and origin calls.

Does this service need a warm-up? A decision on How does the first minute after a restart look with 3 outcomes. Does this service need a warm-up? How does the first minute after a restart look? refills from traffic in seconds no warm-up less startup code shared Redis tier exists warm local from Redis 66 µs per key origin share spikes, p99 jumps warm hot set, limited deadline on readiness Measure the first minute before and after; warm-up is only worth its startup risk if the numbers move.

Verification

Warm-up is done right when:

  • Only a measured hot set blocks readiness, and only up to a deadline.
  • Origin concurrency during warm-up stays within the budget, per worker and per fleet.
  • A slow origin never stops workers starting; they start partially warm.
  • First-minute metrics show a lower origin share and p99 than without warming.

Diagnostic Hook: log warm-up duration, keys warmed and keys failed at every start, and export the first-minute hit ratio per tier. Alert when warm-up hits its deadline on most workers in a deploy — that is an origin that is already struggling, and the deploy is about to add load to it.

Pitfalls & edge cases

  • Unbounded gather over every key. The warm-up becomes the load spike.
  • Readiness blocked on the full key set. Startup time grows with the cache, and an origin outage prevents recovery.
  • Warming keys nobody reads. Old hot sets drift; regenerate them from recent traffic.
  • Warm-up failures aborting startup. Catch per key; a mostly-warm worker is fine.

Frequently Asked Questions

Should I warm my cache before a service accepts traffic?

Warm only the small set of hottest keys, with a deadline, before reporting ready, and warm the rest in the background afterwards. Measure the first minute after a restart to confirm warming lowers origin calls and latency.

How fast should a cache warm-up run?

As fast as the origin can tolerate on top of live traffic. Limit concurrency per worker and divide the budget by the number of workers that start together; an unbounded warm-up of 1,000 keys put 1,000 concurrent requests on the origin in testing.

How do I avoid a load spike when the whole fleet restarts?

Add random jitter before each worker's warm-up, scale per-worker concurrency down with the replica count, and warm from a shared Redis tier when one exists so only the first worker reaches the origin for each key.

What if the origin is down during warm-up?

Bound the warm-up with a timeout and mark the worker ready anyway. A partially warm worker serving traffic is better than a fleet that cannot start.