Skip to content

Isolating Tenants with Bulkheads

In a multi-tenant service, one customer's traffic can consume the capacity everyone shares: a bulk import, a misbehaving integration, a report that hits a slow path. A global concurrency limit protects the service from overload but not tenants from each other — the noisy tenant fills the limit and everyone else queues behind it. Measured with asyncio: one tenant sent 200 requests taking 1 s each while another sent 50 requests taking 20 ms, spread over a second. Behind a shared limit of 20, the quiet tenant's median latency was 9,558 ms and its worst 10,023 ms — it waited behind the noisy tenant's whole backlog. With a limit of 10 per tenant, the quiet tenant's median was 20 ms, its worst 21 ms, and the noisy tenant still completed all 200 requests, just more slowly. Rejecting the noisy tenant's requests beyond its limit instead of queueing them served 10 and rejected 190, with the same 20 ms for the quiet tenant. This guide builds per-tenant bulkheads and chooses between queueing and rejecting.

Prerequisites

1. See how a shared limit couples tenants

A single semaphore around the expensive work is the usual first defence against overload:

shared = asyncio.Semaphore(20)


async def handle(tenant: str, request) -> Response:
    async with shared:                       # every tenant competes for the same 20 slots
        return await expensive_work(request)

Measured: the noisy tenant's 200 one-second requests arrived first and occupied all 20 slots for ten seconds; the quiet tenant's 50 fast requests queued behind them, with a median latency of 9,558 ms for work that takes 20 ms. Nothing failed, which is what makes this hard to spot — the service looks healthy, throughput is at its limit, and one tenant's users experience a ten-second outage. The semaphore's FIFO queue is fair per request, not per tenant.

Verify: chart latency per tenant during a load test where one tenant sends a burst; with a shared limit, every tenant's latency rises together.

Quiet tenant latency while a noisy tenant floods 3 horizontal bars comparing shared limit 20 with the others. Quiet tenant latency while a noisy tenant floods shared limit 20 9,558 ms p50 10 per tenant, queue excess 20 ms p50 10 per tenant, reject excess 20 ms p50 asyncio; noisy tenant 200 requests of 1 s, quiet tenant 50 requests of 20 ms over 1 s. Per-tenant limits decouple tenants; the global limit alone does not.

2. Give each tenant its own limit

A semaphore per tenant caps how much of the shared capacity one tenant can occupy:

import collections


class TenantBulkheads:
    def __init__(self, per_tenant: int, total: int) -> None:
        self.per_tenant = per_tenant
        self.total = asyncio.Semaphore(total)                         # protects the service
        self.tenants: dict[str, asyncio.Semaphore] = collections.defaultdict(
            lambda: asyncio.Semaphore(self.per_tenant))               # protects tenants from each other

    @contextlib.asynccontextmanager
    async def slot(self, tenant: str):
        async with self.tenants[tenant]:                              # tenant limit first
            async with self.total:                                    # then the global limit
                yield

Measured with 10 per tenant: the quiet tenant's median latency stayed at 20 ms and its worst at 21 ms while the noisy tenant worked through its 200 requests ten at a time. Acquiring the tenant's semaphore first matters: a tenant over its own limit waits without holding a global slot. Keep the global limit too — per-tenant limits multiplied by the number of active tenants can exceed what the service can run, and the global limit is what protects the process.

Verify: under a noisy-tenant load test, other tenants' latency stays at its normal level while the noisy tenant's rises.

3. Reject instead of queueing when latency matters

A tenant over its limit can wait or be told no. Waiting turns excess load into latency; rejecting turns it into a fast, explicit error the tenant's client can back off from:

class TenantLimiter:
    def __init__(self, per_tenant: int) -> None:
        self.per_tenant = per_tenant
        self.in_flight: collections.Counter[str] = collections.Counter()

    @contextlib.asynccontextmanager
    async def slot(self, tenant: str):
        if self.in_flight[tenant] >= self.per_tenant:
            raise TenantOverLimit(tenant)                 # becomes 429 with Retry-After
        self.in_flight[tenant] += 1
        try:
            yield
        finally:
            self.in_flight[tenant] -= 1

Measured: the noisy tenant got 10 requests served and 190 rejected immediately, and the quiet tenant was unaffected. Rejection keeps queues — and the memory and timeouts that come with them — from building up, and gives the noisy tenant's client a clear signal. A bounded queue per tenant is the middle ground: a few requests may wait, the rest are rejected, as in rate limiting incoming requests in ASGI apps.

Verify: a tenant exceeding its limit receives 429 responses within milliseconds, with no growth in server memory.

A request through tenant bulkheads A flow of 5 stages. A request through tenant bulkheads identify tenant from auth tenant limit full -> 429 or short wait global limit protects the process do the work per-tenant metrics release both in finally Tenant limit first, global limit second, so waiting never holds shared capacity.

4. Size limits from plans, not just fairness

Equal limits for every tenant are simple; real services usually tie limits to what tenants pay for or need:

LIMITS = {"free": 2, "standard": 10, "enterprise": 50}


def limit_for(tenant: Tenant) -> int:
    return tenant.override_limit or LIMITS[tenant.plan]


class PlannedBulkheads(TenantBulkheads):
    def semaphore(self, tenant: Tenant) -> asyncio.Semaphore:
        sem = self.tenants.get(tenant.id)
        if sem is None:
            sem = self.tenants[tenant.id] = asyncio.Semaphore(limit_for(tenant))
        return sem

The sum of limits for tenants that are typically active at once should fit within the global limit with some headroom; oversubscription is normal (not every tenant is busy at once), and the global limit catches the rare moment when they are. With several processes, each has its own semaphores, so a tenant's effective limit is per-process limit × processes; for exact global limits per tenant, keep counters in Redis, at the cost of a round trip per request. Remove idle tenants' semaphores periodically so the dictionary does not grow with every tenant ever seen.

Verify: the configured limits, multiplied by typical concurrent tenants, fit within the global capacity.

5. Isolate the dependencies tenants share

Request concurrency is only one shared resource. Tenants also share database connections, outbound API quotas and queues — and a noisy tenant can exhaust any of them:

async def tenant_query(tenant: str, sql: str, *args):
    async with db_bulkheads.slot(tenant):                    # per-tenant cap on DB connections
        async with pool.acquire() as conn:
            return await conn.fetch(sql, *args)


async def enqueue_job(tenant: str, job: Job) -> None:
    await queues[tenant].put(job)                            # per-tenant queue, served round-robin

A per-tenant cap on database connections keeps one tenant's slow queries from holding the whole pool; per-tenant job queues served round-robin keep one tenant's backlog from delaying everyone's jobs. The same bulkhead idea applies to every pool a tenant can exhaust. Measure per tenant everywhere — latency, in-flight, rejections — because aggregate metrics hide exactly the problem bulkheads solve.

Verify: a tenant running slow queries cannot hold more than its share of database connections, and other tenants' query latency is unaffected.

How should this service isolate tenants? A decision on What does the tenant's excess load deserve with 4 outcomes. How should this service isolate tenants? What does the tenant's excess load deserve? interactive API per-tenant limit, reject with 429 fast feedback batch / background per-tenant limit, bounded queue completes later paid tiers limits by plan sum fits global cap shared DB, queues, APIs bulkhead each one round-robin service Isolation per tenant, protection per process, both measured per tenant.

Verification

Tenants are isolated when:

  • Each tenant has its own limit, acquired before the global one.
  • Excess load is rejected or bounded, never an unbounded queue.
  • Limits reflect plans and fit the global capacity.
  • Shared dependencies — databases, queues, quotas — have per-tenant bulkheads too.

Diagnostic Hook: export in-flight requests, rejections and latency percentiles labelled by tenant (or tenant tier, if cardinality is a concern). One tenant's in-flight pinned at its limit while others are normal is the bulkhead working; every tenant's latency rising together means a shared resource without a per-tenant limit.

Pitfalls & edge cases

  • A global limit only. Measured: the quiet tenant's p50 rose to 9,558 ms.
  • Global limit acquired before the tenant limit. Waiting tenants hold shared slots.
  • Per-process limits mistaken for global ones. Multiply by the number of processes.
  • Unbounded per-tenant structures. Clean up idle tenants' semaphores.

Frequently Asked Questions

How do I stop one tenant from slowing down others in an asyncio service?

Give each tenant its own concurrency limit, acquired before a global limit, and reject or bound requests beyond it. In testing, a quiet tenant's median latency stayed at 20 ms during another tenant's flood, against 9,558 ms with only a shared limit.

Should excess tenant requests wait or be rejected?

Reject with 429 for interactive APIs, so clients back off and queues do not build; use a small bounded queue for batch work that should eventually complete.

What is a bulkhead in software?

An isolation boundary that limits how much of a shared resource one consumer can use, so a failure or overload in one part cannot sink the rest, named after a ship's watertight compartments.

How do per-tenant limits work across multiple processes?

Each process has its own semaphores, so the effective limit is per-process limit times processes. For exact global per-tenant limits, keep counters in a shared store such as Redis.