Skip to content

Rate Limiting WebSocket Messages per Connection

Request-level rate limiting protects HTTP endpoints, but a WebSocket is one request: once the handshake succeeds, the client can send as many messages as it likes over the open connection, and every one of them runs handler code on your event loop. A per-connection limit on messages closes that gap. Tested with the websockets library, a connection limited to 20 messages per second with a burst of 40 that received a flood of 1,000 messages handled the first 40, dropped the next 100, and was then closed by the server with code 1008 and reason message rate exceeded; a client pacing itself at 15 messages per second had all 60 of its messages handled and none dropped. This guide builds that limiter, decides what to do with excess messages, and extends it to users with several connections.

Prerequisites

1. Give each connection its own bucket

The bucket lives in the handler's scope, so it is created with the connection and discarded with it — no table to clean up:

import time
from websockets.asyncio.server import serve


class TokenBucket:
    def __init__(self, rate: float, burst: int) -> None:
        self.rate, self.burst = rate, burst
        self.tokens, self.updated = float(burst), time.monotonic()

    def allow(self, cost: float = 1.0) -> bool:
        now = time.monotonic()
        self.tokens = min(self.burst, self.tokens + (now - self.updated) * self.rate)
        self.updated = now
        if self.tokens >= cost:
            self.tokens -= cost
            return True
        return False


async def handler(ws):
    bucket = TokenBucket(rate=20, burst=40)
    async for message in ws:
        if not bucket.allow():
            continue                              # step 2 decides what really happens here
        await process(ws, message)


async def main():
    async with serve(handler, "0.0.0.0", 8765) as server:
        await server.serve_forever()

Size the burst for legitimate bursts — a client reconnecting and replaying its subscriptions, a user typing quickly — and the rate for sustained use. Tested: a client sending 15 messages per second against a 20-per-second limit had nothing dropped.

Verify: a paced client below the rate never has messages dropped; a burst beyond burst does.

2. Drop, warn, then close

Silently dropping excess messages hides bugs from honest clients, and never closing lets an abusive one keep the connection busy. A three-stage policy handles both:

async def handler(ws):
    bucket = TokenBucket(rate=20, burst=40)
    strikes = 0
    async for message in ws:
        if bucket.allow():
            await process(ws, message)
            continue
        strikes += 1
        if strikes == 1:
            await ws.send('{"error": "rate_limited", "retry_after_ms": 100}')   # tell the client once
        if strikes >= 100:
            await ws.close(code=1008, reason="message rate exceeded")             # policy violation
            return

Close code 1008 is the WebSocket "policy violation" code, which well-written clients treat as "do not reconnect immediately". Measured on a 1,000-message flood: 40 handled, 100 dropped, then closed with 1008 and the reason string visible to the client. Reset strikes when the client behaves for a while if you want to forgive short bursts; the cumulative count shown here is the strict version.

Verify: a flooding test client receives one rate-limit notice, then a 1008 close after the strike threshold.

A flooding connection, message by message 3 lanes over time. A flooding connection, message by message client sends 1,000 messages as fast as it can server bucket 40 handled 100 dropped close 1008 client sees one rate_limited notice close: rate exceeded time → Measured: 40 handled, 100 dropped, then the connection was closed with a reason the client can read.

3. Charge by message cost, not count

Messages are not equal. A subscription request that fans out to a database query costs far more than a ping. Charge each message type its cost against the same bucket:

import json

COSTS = {"ping": 0.1, "chat": 1.0, "subscribe": 5.0, "history": 10.0}


async def handler(ws):
    bucket = TokenBucket(rate=20, burst=40)
    async for raw in ws:
        if len(raw) > 64_000:
            await ws.close(code=1009, reason="message too big")         # 1009: too large
            return
        msg = json.loads(raw)
        if not bucket.allow(COSTS.get(msg.get("type"), 1.0)):
            ...                                                          # drop / warn / close
            continue
        await dispatch(ws, msg)

Check size before parsing: json.loads on a multi-megabyte frame is CPU on the event loop that the bucket never saw. The websockets server also enforces max_size (1 MiB by default) at the protocol level; set it to the largest legitimate message rather than relying on the default. Per-message processing that is genuinely heavy should be queued, so a burst of expensive messages applies backpressure rather than stalling the loop.

Verify: a client sending only history requests is limited at a fifth of the rate of one sending chat.

What to do with a message over the limit A grid of 4 rows by 3 columns. What to do with a message over the limit action fits risk drop silently lossy telemetry, cursor moves hides client bugs drop and notify once interactive apps client must handle notices close with 1008 sustained abuse honest clients must back off queue for later must-not-lose commands memory grows under abuse Most services combine the middle two: notify on the first excess, close on sustained excess.

4. Limit per user across connections

A user who opens ten connections gets ten buckets. When the limit is about the user rather than the socket — an API quota, a chat spam rule — keep a bucket per user id, shared by all their connections in the process:

from collections import defaultdict

user_buckets: dict[str, TokenBucket] = {}
user_connections: dict[str, int] = defaultdict(int)
MAX_CONNECTIONS = 5


async def handler(ws):
    user = await authenticate(ws)
    if user_connections[user] >= MAX_CONNECTIONS:
        await ws.close(code=1008, reason="too many connections")
        return
    user_connections[user] += 1
    bucket = user_buckets.setdefault(user, TokenBucket(rate=20, burst=40))
    try:
        async for message in ws:
            if bucket.allow():
                await process(ws, message)
    finally:
        user_connections[user] -= 1
        if user_connections[user] == 0:
            user_buckets.pop(user, None)
            del user_connections[user]

Capping connections per user matters as much as capping messages: each open socket costs a task, buffers and a file descriptor. Across processes, the connection count and bucket need a shared store, as in scaling WebSockets across processes with Redis pub/sub. Authentication itself belongs at the handshake, covered in authenticating WebSocket connections.

Verify: a user with five open connections cannot open a sixth, and their total message rate across connections stays at the user limit.

5. Protect the outbound direction too

Rate limits on inbound messages do not stop your handler from flooding a slow client — a chatty subscription can produce messages faster than the client reads them. The outbound side needs its own bound:

async def send_limited(ws, out_bucket: TokenBucket, payload: str) -> bool:
    if not out_bucket.allow():
        return False                                   # coalesce or drop this update
    await ws.send(payload)                             # waits if the client is not reading
    return True

await ws.send() already applies backpressure when the connection's write buffer is full, which protects memory but stalls the sending task; a per-connection outbound limit decides which updates to send under load — typically the latest state rather than every intermediate one. The slow-consumer side of this is in handling WebSocket backpressure with slow consumers.

Verify: a client that stops reading does not cause server memory to grow, and resumes receiving current state when it starts reading again.

Every inbound frame passes three checks A flow of 3 stages. Every inbound frame passes three checks frame size close 1009 if too big charge cost to bucket per connection or user dispatch or warn, then close 1008 Size before parse, cost before work, and a close the client can understand.

Verification

Message limiting is in place when:

  • Each connection has a bucket sized for real bursts, and paced clients are never limited.
  • Excess produces a notice, then a 1008 close on sustained abuse.
  • Message cost and size are checked before any expensive work.
  • Per-user connection counts and rates are bounded where the limit is about the user.

Diagnostic Hook: export per-connection dropped messages and 1008 closes as counters, and the distribution of message rates per connection. A spike in 1008 closes after a client release usually means that release introduced a tight loop; a long tail of connections running at exactly the limit is a client polling over the socket that should be using subscriptions instead.

Pitfalls & edge cases

  • Relying on HTTP rate limits. They stop at the handshake; every subsequent frame is unlimited.
  • Parsing before size checks. Huge frames cost CPU before the limiter sees them.
  • Buckets keyed by user without cleanup. Remove them when the user's last connection closes.
  • Closing with a generic code. Use 1008 with a reason so clients can distinguish abuse from failure.

Frequently Asked Questions

How do I rate limit messages on a WebSocket connection?

Create a token bucket in the connection handler and check it for every received message. Messages over the limit are dropped, the client is notified once, and sustained excess closes the connection with code 1008.

What close code should I use when a WebSocket client sends too many messages?

1008, policy violation, with a short reason such as message rate exceeded. Use 1009 for messages that are too large.

Do HTTP rate limits protect WebSocket endpoints?

Only the handshake. After the connection is upgraded, every message runs on the same connection with no further HTTP request to limit, so messages need their own per-connection or per-user limit.

How do I limit a user who opens many WebSocket connections?

Cap connections per user at the handshake and keep one bucket per user shared by all their connections. Across several processes, keep both in a shared store such as Redis.