Caching HTTP Responses with ETags in Async Clients¶
An async service that polls or repeatedly reads the same upstream resources — a config endpoint, a product catalogue, another team's API — downloads and parses the same bytes over and over. HTTP already has the fix: the server labels each representation with an ETag, the client sends it back as If-None-Match, and if nothing changed the server answers 304 Not Modified with an empty body. Measured against a local aiohttp server returning a 238 KB JSON document, a revalidation returned 0 body bytes and completed in 5.6 ms at the median against 8.2 ms for a full response; with 5 ms of that being the server's own work, the 304 saved about 2.6 ms of transfer and parsing per call, and 196 of 200 polls were 304s. httpx does not cache by itself, so this guide adds a small ETag layer, then covers the HTTP rules that keep it correct.
Prerequisites¶
- Python 3.11+,
pip install httpx; the server side in the examples usesaiohttp. - Client reuse, from Async HTTP Clients & Servers.
- Cache layering, from building a two-tier local and Redis cache.
1. Store the ETag with the parsed body¶
Keep, per URL, the last ETag and the decoded value. Send the tag on the next request, and reuse the stored value on 304:
import httpx
class ETagCache:
def __init__(self, client: httpx.AsyncClient) -> None:
self.client = client
self.store: dict[str, tuple[str, object]] = {}
async def get_json(self, url: str):
headers = {}
cached = self.store.get(url)
if cached is not None:
headers["If-None-Match"] = cached[0]
r = await self.client.get(url, headers=headers)
if r.status_code == 304 and cached is not None:
return cached[1] # unchanged: reuse the parsed value
r.raise_for_status()
data = r.json()
if etag := r.headers.get("ETag"):
self.store[url] = (etag, data)
return data
Storing the parsed value, not the raw bytes, is where most of the client-side saving comes from: a 304 skips both the download and json.loads() of 238 KB. The server still does whatever work it needs to decide whether the resource changed — 5 ms in the test — so the saving is transfer plus parsing, not the whole request.
Verify: after the first call, repeat calls return 304 with an empty body and the same parsed object.
2. Respect Cache-Control before revalidating¶
ETags answer "has it changed?"; Cache-Control answers "do I even need to ask?". A response with max-age=60 can be reused for 60 seconds without any request at all; no-cache means "always revalidate"; no-store means "never keep it". Honour them:
import time
def max_age(headers: httpx.Headers) -> float:
cc = headers.get("Cache-Control", "")
if "no-store" in cc or "no-cache" in cc:
return 0.0
for part in cc.split(","):
k, _, v = part.strip().partition("=")
if k == "max-age" and v.isdigit():
return float(v)
return 0.0
class FreshnessCache(ETagCache):
async def get_json(self, url: str):
cached = self.store.get(url)
if cached is not None and len(cached) == 3 and cached[2] > time.monotonic():
return cached[1] # still fresh: no request at all
data = await super().get_json(url)
... # store expiry from max_age(response.headers)
return data
A fresh response costs nothing; a revalidated one costs a round trip with an empty body; a changed one costs a full download. Most APIs that send ETags for rapidly changing data also send no-cache, so revalidation is the common case — the test server here did exactly that. For production use, the hishel package implements RFC 9111 caching as an httpx transport, including Vary and storage backends, and is worth using instead of growing this class indefinitely.
Verify: a max-age=60 response produces no requests for 60 seconds; a no-store response is never stored.
3. Collapse concurrent revalidations¶
In an async service, a hundred requests can arrive for the same upstream URL at once, each deciding to revalidate. Share one in-flight request per URL:
import asyncio
class CollapsingETagCache(ETagCache):
def __init__(self, client: httpx.AsyncClient) -> None:
super().__init__(client)
self._inflight: dict[str, asyncio.Task] = {}
async def get_json(self, url: str):
task = self._inflight.get(url)
if task is None:
task = asyncio.ensure_future(super().get_json(url))
self._inflight[url] = task
task.add_done_callback(lambda _: self._inflight.pop(url, None))
return await asyncio.shield(task)
This is the single-flight pattern from implementing the single-flight pattern for duplicate calls, applied at the HTTP layer: a burst of 100 callers produces one conditional GET, and they all get the same parsed value. Combined with a short max-age honoured locally, a hot upstream resource sees at most one request per freshness window per worker.
Verify: under 100 concurrent callers for one URL, the server logs one request.
4. Handle Vary, auth and per-user responses¶
A cache keyed only by URL is wrong when the response depends on request headers. The server signals this with Vary: Vary: Accept-Language means the French and English responses are different representations with different ETags. Include varying headers in the key:
def cache_key(url: str, request_headers: dict[str, str], vary: str | None) -> tuple:
if not vary:
return (url,)
names = sorted(h.strip().lower() for h in vary.split(","))
return (url, *((n, request_headers.get(n, "")) for n in names))
Treat Authorization with particular care. Responses for different users can share a URL and an ETag scheme while containing different data; a client-side cache shared across users must key on the user (or skip caching authenticated responses). A cache inside one service that calls an upstream with one service credential is fine; a cache in a gateway that forwards user tokens is not, unless the key includes the user.
Verify: two requests with different Accept-Language values are cached separately and never served to each other.
5. Bound the store and expire stale entries¶
The ETag store holds parsed bodies, which can be large. Bound it like any other in-memory cache, and drop entries that have not been used for a while — an entry whose ETag is stale still costs a full download on the next change, so keeping it forever buys nothing:
from collections import OrderedDict
class BoundedETagStore:
def __init__(self, maxsize: int = 512) -> None:
self._d: OrderedDict[str, tuple[str, object]] = OrderedDict()
self.maxsize = maxsize
def get(self, url: str):
item = self._d.get(url)
if item is not None:
self._d.move_to_end(url)
return item
def put(self, url: str, etag: str, value) -> None:
self._d[url] = (etag, value)
self._d.move_to_end(url)
while len(self._d) > self.maxsize:
self._d.popitem(last=False)
Size the bound by memory, not entry count, if bodies vary widely — the techniques are in bounding in-memory async caches by size.
Verify: under a workload touching many URLs, the store stays at maxsize and the 304 rate for hot URLs is unchanged.
Verification¶
Client-side HTTP caching works when:
- Unchanged resources return 304 with empty bodies, and the stored value is reused.
Cache-Controlis honoured: fresh responses make no request,no-storeis never stored.- Concurrent callers share one revalidation per URL.
Varyand per-user responses produce separate cache entries.
Diagnostic Hook: count responses per URL by outcome — fresh hit, 304, 200 — and the bytes downloaded. A URL with a high 200 rate despite a stable payload usually means the server's ETag changes on every response (a timestamp or random value in the hash), which is worth reporting upstream; a 304 rate near 100% with no-cache suggests the server could safely send a short max-age and save the round trips entirely.
Pitfalls & edge cases¶
- Weak ETags (
W/"..."). Valid forIf-None-Match; send them back exactly as received. - Returning a shared mutable object. Callers that mutate the cached dict corrupt every later 304; return copies or immutable structures.
- Ignoring
Vary. One user's language, encoding or permissions served to another. - Caching error responses. Only store successful responses that carry a validator.
Frequently Asked Questions¶
Does httpx cache responses automatically?
No. httpx sends exactly the requests you make. Add If-None-Match yourself as shown, or use a caching transport such as hishel, which implements HTTP caching rules for httpx.
How do ETags reduce bandwidth?
The client sends the stored ETag in If-None-Match; if the resource is unchanged the server replies 304 Not Modified with an empty body. In testing a 238 KB JSON document revalidated with zero body bytes.
What is the difference between max-age and an ETag?
max-age says how long a response may be reused without asking the server at all. An ETag lets the client ask cheaply whether its copy is still valid. Fresh responses cost nothing; revalidations cost a round trip with no body.
Is it safe to cache authenticated API responses on the client?
Only if the cache key includes the user or credential, or the cache is used by a single identity. Otherwise one user's response can be served to another.
Related¶
- Async Caching & Deduplication — up to the topic overview.
- Retrying httpx requests with transport retries — another transport-level concern for the same clients.
- Concurrent Execution & Worker Patterns — the section overview.