Skip to content

Caching HTTP Responses with ETags in Async Clients

An async service that polls or repeatedly reads the same upstream resources — a config endpoint, a product catalogue, another team's API — downloads and parses the same bytes over and over. HTTP already has the fix: the server labels each representation with an ETag, the client sends it back as If-None-Match, and if nothing changed the server answers 304 Not Modified with an empty body. Measured against a local aiohttp server returning a 238 KB JSON document, a revalidation returned 0 body bytes and completed in 5.6 ms at the median against 8.2 ms for a full response; with 5 ms of that being the server's own work, the 304 saved about 2.6 ms of transfer and parsing per call, and 196 of 200 polls were 304s. httpx does not cache by itself, so this guide adds a small ETag layer, then covers the HTTP rules that keep it correct.

Prerequisites

1. Store the ETag with the parsed body

Keep, per URL, the last ETag and the decoded value. Send the tag on the next request, and reuse the stored value on 304:

import httpx


class ETagCache:
    def __init__(self, client: httpx.AsyncClient) -> None:
        self.client = client
        self.store: dict[str, tuple[str, object]] = {}

    async def get_json(self, url: str):
        headers = {}
        cached = self.store.get(url)
        if cached is not None:
            headers["If-None-Match"] = cached[0]
        r = await self.client.get(url, headers=headers)
        if r.status_code == 304 and cached is not None:
            return cached[1]                              # unchanged: reuse the parsed value
        r.raise_for_status()
        data = r.json()
        if etag := r.headers.get("ETag"):
            self.store[url] = (etag, data)
        return data

Storing the parsed value, not the raw bytes, is where most of the client-side saving comes from: a 304 skips both the download and json.loads() of 238 KB. The server still does whatever work it needs to decide whether the resource changed — 5 ms in the test — so the saving is transfer plus parsing, not the whole request.

Verify: after the first call, repeat calls return 304 with an empty body and the same parsed object.

A conditional GET that returns nothing A sequence of 6 messages between 3 participants. A conditional GET that returns nothing client cache httpx server GET /items 200, ETag: "9f2c", 238 KB store ETag + parsed JSON GET /items, If-None-Match: "9f2c" 304 Not Modified, 0 bytes return stored value The server still decides; the client just stops downloading and parsing what it already has.

2. Respect Cache-Control before revalidating

ETags answer "has it changed?"; Cache-Control answers "do I even need to ask?". A response with max-age=60 can be reused for 60 seconds without any request at all; no-cache means "always revalidate"; no-store means "never keep it". Honour them:

import time


def max_age(headers: httpx.Headers) -> float:
    cc = headers.get("Cache-Control", "")
    if "no-store" in cc or "no-cache" in cc:
        return 0.0
    for part in cc.split(","):
        k, _, v = part.strip().partition("=")
        if k == "max-age" and v.isdigit():
            return float(v)
    return 0.0


class FreshnessCache(ETagCache):
    async def get_json(self, url: str):
        cached = self.store.get(url)
        if cached is not None and len(cached) == 3 and cached[2] > time.monotonic():
            return cached[1]                              # still fresh: no request at all
        data = await super().get_json(url)
        ...                                               # store expiry from max_age(response.headers)
        return data

A fresh response costs nothing; a revalidated one costs a round trip with an empty body; a changed one costs a full download. Most APIs that send ETags for rapidly changing data also send no-cache, so revalidation is the common case — the test server here did exactly that. For production use, the hishel package implements RFC 9111 caching as an httpx transport, including Vary and storage backends, and is worth using instead of growing this class indefinitely.

Verify: a max-age=60 response produces no requests for 60 seconds; a no-store response is never stored.

3. Collapse concurrent revalidations

In an async service, a hundred requests can arrive for the same upstream URL at once, each deciding to revalidate. Share one in-flight request per URL:

import asyncio


class CollapsingETagCache(ETagCache):
    def __init__(self, client: httpx.AsyncClient) -> None:
        super().__init__(client)
        self._inflight: dict[str, asyncio.Task] = {}

    async def get_json(self, url: str):
        task = self._inflight.get(url)
        if task is None:
            task = asyncio.ensure_future(super().get_json(url))
            self._inflight[url] = task
            task.add_done_callback(lambda _: self._inflight.pop(url, None))
        return await asyncio.shield(task)

This is the single-flight pattern from implementing the single-flight pattern for duplicate calls, applied at the HTTP layer: a burst of 100 callers produces one conditional GET, and they all get the same parsed value. Combined with a short max-age honoured locally, a hot upstream resource sees at most one request per freshness window per worker.

Verify: under 100 concurrent callers for one URL, the server logs one request.

Full response versus revalidation, same 238 KB resource 2 horizontal bars comparing 200 with body with the others. Full response versus revalidation, same 238 KB resource 200 with body 8.2 ms, 237,780 bytes 304 Not Modified 5.6 ms, 0 bytes Local aiohttp server doing 5 ms of work per request; httpx client; 196 of 200 polls were 304s. Revalidation removes transfer and parsing; the server's own work remains.

4. Handle Vary, auth and per-user responses

A cache keyed only by URL is wrong when the response depends on request headers. The server signals this with Vary: Vary: Accept-Language means the French and English responses are different representations with different ETags. Include varying headers in the key:

def cache_key(url: str, request_headers: dict[str, str], vary: str | None) -> tuple:
    if not vary:
        return (url,)
    names = sorted(h.strip().lower() for h in vary.split(","))
    return (url, *((n, request_headers.get(n, "")) for n in names))

Treat Authorization with particular care. Responses for different users can share a URL and an ETag scheme while containing different data; a client-side cache shared across users must key on the user (or skip caching authenticated responses). A cache inside one service that calls an upstream with one service credential is fine; a cache in a gateway that forwards user tokens is not, unless the key includes the user.

Verify: two requests with different Accept-Language values are cached separately and never served to each other.

5. Bound the store and expire stale entries

The ETag store holds parsed bodies, which can be large. Bound it like any other in-memory cache, and drop entries that have not been used for a while — an entry whose ETag is stale still costs a full download on the next change, so keeping it forever buys nothing:

from collections import OrderedDict


class BoundedETagStore:
    def __init__(self, maxsize: int = 512) -> None:
        self._d: OrderedDict[str, tuple[str, object]] = OrderedDict()
        self.maxsize = maxsize

    def get(self, url: str):
        item = self._d.get(url)
        if item is not None:
            self._d.move_to_end(url)
        return item

    def put(self, url: str, etag: str, value) -> None:
        self._d[url] = (etag, value)
        self._d.move_to_end(url)
        while len(self._d) > self.maxsize:
            self._d.popitem(last=False)

Size the bound by memory, not entry count, if bodies vary widely — the techniques are in bounding in-memory async caches by size.

Verify: under a workload touching many URLs, the store stays at maxsize and the 304 rate for hot URLs is unchanged.

What should the client do with this response? A decision on What did Cache-Control say with 3 outcomes. What should the client do with this response? What did Cache-Control say? max-age still fresh reuse, no request free no-cache, or stale with ETag conditional GET 304 is 0 bytes no-store never cache always full fetch Freshness avoids requests; validators avoid bodies; no-store avoids caching entirely.

Verification

Client-side HTTP caching works when:

  • Unchanged resources return 304 with empty bodies, and the stored value is reused.
  • Cache-Control is honoured: fresh responses make no request, no-store is never stored.
  • Concurrent callers share one revalidation per URL.
  • Vary and per-user responses produce separate cache entries.

Diagnostic Hook: count responses per URL by outcome — fresh hit, 304, 200 — and the bytes downloaded. A URL with a high 200 rate despite a stable payload usually means the server's ETag changes on every response (a timestamp or random value in the hash), which is worth reporting upstream; a 304 rate near 100% with no-cache suggests the server could safely send a short max-age and save the round trips entirely.

Pitfalls & edge cases

  • Weak ETags (W/"..."). Valid for If-None-Match; send them back exactly as received.
  • Returning a shared mutable object. Callers that mutate the cached dict corrupt every later 304; return copies or immutable structures.
  • Ignoring Vary. One user's language, encoding or permissions served to another.
  • Caching error responses. Only store successful responses that carry a validator.

Frequently Asked Questions

Does httpx cache responses automatically?

No. httpx sends exactly the requests you make. Add If-None-Match yourself as shown, or use a caching transport such as hishel, which implements HTTP caching rules for httpx.

How do ETags reduce bandwidth?

The client sends the stored ETag in If-None-Match; if the resource is unchanged the server replies 304 Not Modified with an empty body. In testing a 238 KB JSON document revalidated with zero body bytes.

What is the difference between max-age and an ETag?

max-age says how long a response may be reused without asking the server at all. An ETag lets the client ask cheaply whether its copy is still valid. Fresh responses cost nothing; revalidations cost a round trip with no body.

Is it safe to cache authenticated API responses on the client?

Only if the cache key includes the user or credential, or the cache is used by a single identity. Otherwise one user's response can be served to another.