Rate Limiting WebSocket Messages per Connection¶
Request-level rate limiting protects HTTP endpoints, but a WebSocket is one request: once the handshake succeeds, the client can send as many messages as it likes over the open connection, and every one of them runs handler code on your event loop. A per-connection limit on messages closes that gap. Tested with the websockets library, a connection limited to 20 messages per second with a burst of 40 that received a flood of 1,000 messages handled the first 40, dropped the next 100, and was then closed by the server with code 1008 and reason message rate exceeded; a client pacing itself at 15 messages per second had all 60 of its messages handled and none dropped. This guide builds that limiter, decides what to do with excess messages, and extends it to users with several connections.
Prerequisites¶
- Python 3.11+,
pip install websockets(the examples use thewebsockets.asyncioAPI of version 13 and later); the same structure applies to Starlette/FastAPI websockets. - Token buckets, from rate limiting incoming requests in ASGI apps.
- WebSocket backpressure, from handling WebSocket backpressure with slow consumers.
1. Give each connection its own bucket¶
The bucket lives in the handler's scope, so it is created with the connection and discarded with it — no table to clean up:
import time
from websockets.asyncio.server import serve
class TokenBucket:
def __init__(self, rate: float, burst: int) -> None:
self.rate, self.burst = rate, burst
self.tokens, self.updated = float(burst), time.monotonic()
def allow(self, cost: float = 1.0) -> bool:
now = time.monotonic()
self.tokens = min(self.burst, self.tokens + (now - self.updated) * self.rate)
self.updated = now
if self.tokens >= cost:
self.tokens -= cost
return True
return False
async def handler(ws):
bucket = TokenBucket(rate=20, burst=40)
async for message in ws:
if not bucket.allow():
continue # step 2 decides what really happens here
await process(ws, message)
async def main():
async with serve(handler, "0.0.0.0", 8765) as server:
await server.serve_forever()
Size the burst for legitimate bursts — a client reconnecting and replaying its subscriptions, a user typing quickly — and the rate for sustained use. Tested: a client sending 15 messages per second against a 20-per-second limit had nothing dropped.
Verify: a paced client below the rate never has messages dropped; a burst beyond burst does.
2. Drop, warn, then close¶
Silently dropping excess messages hides bugs from honest clients, and never closing lets an abusive one keep the connection busy. A three-stage policy handles both:
async def handler(ws):
bucket = TokenBucket(rate=20, burst=40)
strikes = 0
async for message in ws:
if bucket.allow():
await process(ws, message)
continue
strikes += 1
if strikes == 1:
await ws.send('{"error": "rate_limited", "retry_after_ms": 100}') # tell the client once
if strikes >= 100:
await ws.close(code=1008, reason="message rate exceeded") # policy violation
return
Close code 1008 is the WebSocket "policy violation" code, which well-written clients treat as "do not reconnect immediately". Measured on a 1,000-message flood: 40 handled, 100 dropped, then closed with 1008 and the reason string visible to the client. Reset strikes when the client behaves for a while if you want to forgive short bursts; the cumulative count shown here is the strict version.
Verify: a flooding test client receives one rate-limit notice, then a 1008 close after the strike threshold.
3. Charge by message cost, not count¶
Messages are not equal. A subscription request that fans out to a database query costs far more than a ping. Charge each message type its cost against the same bucket:
import json
COSTS = {"ping": 0.1, "chat": 1.0, "subscribe": 5.0, "history": 10.0}
async def handler(ws):
bucket = TokenBucket(rate=20, burst=40)
async for raw in ws:
if len(raw) > 64_000:
await ws.close(code=1009, reason="message too big") # 1009: too large
return
msg = json.loads(raw)
if not bucket.allow(COSTS.get(msg.get("type"), 1.0)):
... # drop / warn / close
continue
await dispatch(ws, msg)
Check size before parsing: json.loads on a multi-megabyte frame is CPU on the event loop that the bucket never saw. The websockets server also enforces max_size (1 MiB by default) at the protocol level; set it to the largest legitimate message rather than relying on the default. Per-message processing that is genuinely heavy should be queued, so a burst of expensive messages applies backpressure rather than stalling the loop.
Verify: a client sending only history requests is limited at a fifth of the rate of one sending chat.
4. Limit per user across connections¶
A user who opens ten connections gets ten buckets. When the limit is about the user rather than the socket — an API quota, a chat spam rule — keep a bucket per user id, shared by all their connections in the process:
from collections import defaultdict
user_buckets: dict[str, TokenBucket] = {}
user_connections: dict[str, int] = defaultdict(int)
MAX_CONNECTIONS = 5
async def handler(ws):
user = await authenticate(ws)
if user_connections[user] >= MAX_CONNECTIONS:
await ws.close(code=1008, reason="too many connections")
return
user_connections[user] += 1
bucket = user_buckets.setdefault(user, TokenBucket(rate=20, burst=40))
try:
async for message in ws:
if bucket.allow():
await process(ws, message)
finally:
user_connections[user] -= 1
if user_connections[user] == 0:
user_buckets.pop(user, None)
del user_connections[user]
Capping connections per user matters as much as capping messages: each open socket costs a task, buffers and a file descriptor. Across processes, the connection count and bucket need a shared store, as in scaling WebSockets across processes with Redis pub/sub. Authentication itself belongs at the handshake, covered in authenticating WebSocket connections.
Verify: a user with five open connections cannot open a sixth, and their total message rate across connections stays at the user limit.
5. Protect the outbound direction too¶
Rate limits on inbound messages do not stop your handler from flooding a slow client — a chatty subscription can produce messages faster than the client reads them. The outbound side needs its own bound:
async def send_limited(ws, out_bucket: TokenBucket, payload: str) -> bool:
if not out_bucket.allow():
return False # coalesce or drop this update
await ws.send(payload) # waits if the client is not reading
return True
await ws.send() already applies backpressure when the connection's write buffer is full, which protects memory but stalls the sending task; a per-connection outbound limit decides which updates to send under load — typically the latest state rather than every intermediate one. The slow-consumer side of this is in handling WebSocket backpressure with slow consumers.
Verify: a client that stops reading does not cause server memory to grow, and resumes receiving current state when it starts reading again.
Verification¶
Message limiting is in place when:
- Each connection has a bucket sized for real bursts, and paced clients are never limited.
- Excess produces a notice, then a 1008 close on sustained abuse.
- Message cost and size are checked before any expensive work.
- Per-user connection counts and rates are bounded where the limit is about the user.
Diagnostic Hook: export per-connection dropped messages and 1008 closes as counters, and the distribution of message rates per connection. A spike in 1008 closes after a client release usually means that release introduced a tight loop; a long tail of connections running at exactly the limit is a client polling over the socket that should be using subscriptions instead.
Pitfalls & edge cases¶
- Relying on HTTP rate limits. They stop at the handshake; every subsequent frame is unlimited.
- Parsing before size checks. Huge frames cost CPU before the limiter sees them.
- Buckets keyed by user without cleanup. Remove them when the user's last connection closes.
- Closing with a generic code. Use 1008 with a reason so clients can distinguish abuse from failure.
Frequently Asked Questions¶
How do I rate limit messages on a WebSocket connection?
Create a token bucket in the connection handler and check it for every received message. Messages over the limit are dropped, the client is notified once, and sustained excess closes the connection with code 1008.
What close code should I use when a WebSocket client sends too many messages?
1008, policy violation, with a short reason such as message rate exceeded. Use 1009 for messages that are too large.
Do HTTP rate limits protect WebSocket endpoints?
Only the handshake. After the connection is upgraded, every message runs on the same connection with no further HTTP request to limit, so messages need their own per-connection or per-user limit.
How do I limit a user who opens many WebSocket connections?
Cap connections per user at the handshake and keep one bucket per user shared by all their connections. Across several processes, keep both in a shared store such as Redis.
Related¶
- Rate Limiting & Throttling — up to the topic overview.
- Broadcasting to thousands of WebSocket clients — the outbound side at scale.
- Concurrent Execution & Worker Patterns — the section overview.