Isolating Tenants with Bulkheads¶
In a multi-tenant service, one customer's traffic can consume the capacity everyone shares: a bulk import, a misbehaving integration, a report that hits a slow path. A global concurrency limit protects the service from overload but not tenants from each other — the noisy tenant fills the limit and everyone else queues behind it. Measured with asyncio: one tenant sent 200 requests taking 1 s each while another sent 50 requests taking 20 ms, spread over a second. Behind a shared limit of 20, the quiet tenant's median latency was 9,558 ms and its worst 10,023 ms — it waited behind the noisy tenant's whole backlog. With a limit of 10 per tenant, the quiet tenant's median was 20 ms, its worst 21 ms, and the noisy tenant still completed all 200 requests, just more slowly. Rejecting the noisy tenant's requests beyond its limit instead of queueing them served 10 and rejected 190, with the same 20 ms for the quiet tenant. This guide builds per-tenant bulkheads and chooses between queueing and rejecting.
Prerequisites¶
- Python 3.11+, stdlib only.
- Per-dependency bulkheads, from bulkhead isolation with per-dependency semaphores.
- Fair scheduling, from fair scheduling across tenants in a worker pool.
1. See how a shared limit couples tenants¶
A single semaphore around the expensive work is the usual first defence against overload:
shared = asyncio.Semaphore(20)
async def handle(tenant: str, request) -> Response:
async with shared: # every tenant competes for the same 20 slots
return await expensive_work(request)
Measured: the noisy tenant's 200 one-second requests arrived first and occupied all 20 slots for ten seconds; the quiet tenant's 50 fast requests queued behind them, with a median latency of 9,558 ms for work that takes 20 ms. Nothing failed, which is what makes this hard to spot — the service looks healthy, throughput is at its limit, and one tenant's users experience a ten-second outage. The semaphore's FIFO queue is fair per request, not per tenant.
Verify: chart latency per tenant during a load test where one tenant sends a burst; with a shared limit, every tenant's latency rises together.
2. Give each tenant its own limit¶
A semaphore per tenant caps how much of the shared capacity one tenant can occupy:
import collections
class TenantBulkheads:
def __init__(self, per_tenant: int, total: int) -> None:
self.per_tenant = per_tenant
self.total = asyncio.Semaphore(total) # protects the service
self.tenants: dict[str, asyncio.Semaphore] = collections.defaultdict(
lambda: asyncio.Semaphore(self.per_tenant)) # protects tenants from each other
@contextlib.asynccontextmanager
async def slot(self, tenant: str):
async with self.tenants[tenant]: # tenant limit first
async with self.total: # then the global limit
yield
Measured with 10 per tenant: the quiet tenant's median latency stayed at 20 ms and its worst at 21 ms while the noisy tenant worked through its 200 requests ten at a time. Acquiring the tenant's semaphore first matters: a tenant over its own limit waits without holding a global slot. Keep the global limit too — per-tenant limits multiplied by the number of active tenants can exceed what the service can run, and the global limit is what protects the process.
Verify: under a noisy-tenant load test, other tenants' latency stays at its normal level while the noisy tenant's rises.
3. Reject instead of queueing when latency matters¶
A tenant over its limit can wait or be told no. Waiting turns excess load into latency; rejecting turns it into a fast, explicit error the tenant's client can back off from:
class TenantLimiter:
def __init__(self, per_tenant: int) -> None:
self.per_tenant = per_tenant
self.in_flight: collections.Counter[str] = collections.Counter()
@contextlib.asynccontextmanager
async def slot(self, tenant: str):
if self.in_flight[tenant] >= self.per_tenant:
raise TenantOverLimit(tenant) # becomes 429 with Retry-After
self.in_flight[tenant] += 1
try:
yield
finally:
self.in_flight[tenant] -= 1
Measured: the noisy tenant got 10 requests served and 190 rejected immediately, and the quiet tenant was unaffected. Rejection keeps queues — and the memory and timeouts that come with them — from building up, and gives the noisy tenant's client a clear signal. A bounded queue per tenant is the middle ground: a few requests may wait, the rest are rejected, as in rate limiting incoming requests in ASGI apps.
Verify: a tenant exceeding its limit receives 429 responses within milliseconds, with no growth in server memory.
4. Size limits from plans, not just fairness¶
Equal limits for every tenant are simple; real services usually tie limits to what tenants pay for or need:
LIMITS = {"free": 2, "standard": 10, "enterprise": 50}
def limit_for(tenant: Tenant) -> int:
return tenant.override_limit or LIMITS[tenant.plan]
class PlannedBulkheads(TenantBulkheads):
def semaphore(self, tenant: Tenant) -> asyncio.Semaphore:
sem = self.tenants.get(tenant.id)
if sem is None:
sem = self.tenants[tenant.id] = asyncio.Semaphore(limit_for(tenant))
return sem
The sum of limits for tenants that are typically active at once should fit within the global limit with some headroom; oversubscription is normal (not every tenant is busy at once), and the global limit catches the rare moment when they are. With several processes, each has its own semaphores, so a tenant's effective limit is per-process limit × processes; for exact global limits per tenant, keep counters in Redis, at the cost of a round trip per request. Remove idle tenants' semaphores periodically so the dictionary does not grow with every tenant ever seen.
Verify: the configured limits, multiplied by typical concurrent tenants, fit within the global capacity.
5. Isolate the dependencies tenants share¶
Request concurrency is only one shared resource. Tenants also share database connections, outbound API quotas and queues — and a noisy tenant can exhaust any of them:
async def tenant_query(tenant: str, sql: str, *args):
async with db_bulkheads.slot(tenant): # per-tenant cap on DB connections
async with pool.acquire() as conn:
return await conn.fetch(sql, *args)
async def enqueue_job(tenant: str, job: Job) -> None:
await queues[tenant].put(job) # per-tenant queue, served round-robin
A per-tenant cap on database connections keeps one tenant's slow queries from holding the whole pool; per-tenant job queues served round-robin keep one tenant's backlog from delaying everyone's jobs. The same bulkhead idea applies to every pool a tenant can exhaust. Measure per tenant everywhere — latency, in-flight, rejections — because aggregate metrics hide exactly the problem bulkheads solve.
Verify: a tenant running slow queries cannot hold more than its share of database connections, and other tenants' query latency is unaffected.
Verification¶
Tenants are isolated when:
- Each tenant has its own limit, acquired before the global one.
- Excess load is rejected or bounded, never an unbounded queue.
- Limits reflect plans and fit the global capacity.
- Shared dependencies — databases, queues, quotas — have per-tenant bulkheads too.
Diagnostic Hook: export in-flight requests, rejections and latency percentiles labelled by tenant (or tenant tier, if cardinality is a concern). One tenant's in-flight pinned at its limit while others are normal is the bulkhead working; every tenant's latency rising together means a shared resource without a per-tenant limit.
Pitfalls & edge cases¶
- A global limit only. Measured: the quiet tenant's p50 rose to 9,558 ms.
- Global limit acquired before the tenant limit. Waiting tenants hold shared slots.
- Per-process limits mistaken for global ones. Multiply by the number of processes.
- Unbounded per-tenant structures. Clean up idle tenants' semaphores.
Frequently Asked Questions¶
How do I stop one tenant from slowing down others in an asyncio service?
Give each tenant its own concurrency limit, acquired before a global limit, and reject or bound requests beyond it. In testing, a quiet tenant's median latency stayed at 20 ms during another tenant's flood, against 9,558 ms with only a shared limit.
Should excess tenant requests wait or be rejected?
Reject with 429 for interactive APIs, so clients back off and queues do not build; use a small bounded queue for batch work that should eventually complete.
What is a bulkhead in software?
An isolation boundary that limits how much of a shared resource one consumer can use, so a failure or overload in one part cannot sink the rest, named after a ship's watertight compartments.
How do per-tenant limits work across multiple processes?
Each process has its own semaphores, so the effective limit is per-process limit times processes. For exact global per-tenant limits, keep counters in a shared store such as Redis.
Related¶
- Circuit Breakers & Bulkheads — up to the topic overview.
- Tuning circuit breaker thresholds from metrics — the other half of protecting shared dependencies.
- Resilience, Cancellation & Error Handling — the section overview.