Capacity Planning with Little's Law¶
Little's law says that the average number of requests inside a system equals the arrival rate times the average time each spends there: L = λ × W. It needs no assumptions about distributions, and for an asyncio service it turns three easy measurements — request rate, latency and the time each request holds a scarce resource — into the numbers capacity planning needs: how many requests are in flight, how many pool connections are busy, and the rate at which the service will fall over. Measured with an aiohttp service on one core whose handler did 10 ms of other I/O and then a 10 ms "query" through a pool of 10 slots, under an open-loop load: at rates from 200 to 900 req/s, λ × W predicted the measured in-flight count within 3% — 4.4 against 4.3 at 200 req/s, 21.3 against 21.3 at 900. Busy pool slots followed λ × 10.6 ms, the query's real holding time. Latency stayed at 21.7–21.9 ms up to 800 req/s with 8.5 of 10 slots busy, reached 23.7 ms at 900, then jumped to 247 ms at 950 and 521 ms at 1,000 as the pool saturated. Doubling the pool to 20 moved the knee to between 1,700 and 1,800 req/s, as the law predicts. This guide uses those relationships to size pools, instances and limits.
Prerequisites¶
- A service you can load at a controlled rate, from building an open-loop load generator in asyncio.
- Request and resource metrics: rate, latency, and time spent holding each pool.
- The topic overview, Load Testing & Benchmarking.
1. Check the law against your own service¶
Little's law holds for any stable system, so it doubles as a check that your metrics are consistent. Measure the arrival rate and mean latency from the client, sample the server's in-flight count, and compare:
in_flight = 0
async def handler(request):
global in_flight
in_flight += 1
try:
return await handle(request)
finally:
in_flight -= 1
async def sample_in_flight(samples: list[int]):
while True:
await asyncio.sleep(0.01)
samples.append(in_flight) # average of samples = L
Measured below saturation: at 200 req/s, mean latency 21.9 ms, λ × W = 4.4, sampled L = 4.3; at 600 req/s, 13.1 against 12.8; at 900 req/s, 21.3 against 21.3. Past saturation the two diverge — at 1,000 req/s, λ × W was 521 and the sampled average 447 — because the queue was growing throughout the run, and the law describes systems in steady state. A system whose measured L keeps rising at a constant arrival rate is not in steady state: it is falling behind.
Verify: at a few rates below saturation, λ × W and the measured in-flight average agree within a few per cent.
2. Apply the law to each scarce resource¶
The law applies to any part of the system, including one pool. The number of busy connections is the request rate times the time each request holds one:
def busy_slots(rate_per_s: float, hold_s: float) -> float:
return rate_per_s * hold_s # L = lambda x W, for the pool
def pool_limit(slots: int, hold_s: float) -> float:
return slots / hold_s # the rate at which every slot is busy
busy_slots(800, 0.0106) # 8.5
pool_limit(10, 0.0106) # ~943 req/s
Measured: 2.1, 4.3, 6.4 and 8.5 busy slots at 200, 400, 600 and 800 req/s — 10.6 ms per request, not the nominal 10 ms, because asyncio.sleep(0.010) and the pool handoff took slightly longer. That small difference moved the predicted limit from 1,000 to about 943 req/s, and the service duly broke between 900 and 950. Measure the real holding time — from acquiring the connection to releasing it, including any work done while holding it — rather than the query's own duration; it is usually longer, and the limit depends on it directly.
Verify: your pool's measured busy count matches rate × holding time, and you know the rate at which it reaches the pool size.
3. Plan for utilisation well below 100%¶
The limit from step 2 is where the system fails, not where it should run. Queueing delay grows sharply as utilisation approaches 100%, because random arrivals bunch up and must wait:
def slots_for(rate_per_s: float, hold_s: float, max_utilisation: float = 0.7) -> int:
return math.ceil(rate_per_s * hold_s / max_utilisation)
slots_for(800, 0.0106) # 13 slots to serve 800 req/s at 65% busy
Measured with 10 slots: at 64% busy (600 req/s) pool waits were zero at p99; at 85% (800 req/s) the p99 wait was 1.3 ms; at 95% (900 req/s) 11.5 ms, and overall p99 latency rose from 25 to 34 ms; at 950 req/s the queue grew without bound and p99 was 473 ms. A planning target of 60–75% utilisation keeps the service on the flat part of that curve and leaves room for bursts and for one instance failing. Doubling the pool to 20 slots was measured too: latency stayed at 21.7–22.7 ms up to 1,700 req/s with 18 slots busy, and broke at 1,800 — the predicted 20 / 10.6 ms ≈ 1,890, less the same margin as before.
Verify: at your planned peak, every pool and worker set is predicted to run below about 75% busy.
4. Turn the numbers into instance counts and limits¶
With per-resource limits known, capacity planning is arithmetic. Work from the peak rate you must serve to the resources each instance needs, then check shared resources across all instances:
peak = 3_000 # req/s to serve at peak
hold_db = 0.0106 # s each request holds a DB connection
per_instance_slots = 20 # pool size per instance
target = 0.7 # max utilisation
per_instance_rate = per_instance_slots / hold_db * target # ~1,320 req/s
instances = math.ceil(peak / per_instance_rate) + 1 # 3 + 1 spare = 4
db_connections = instances * per_instance_slots # 80, must be under max_connections
in_flight_per_instance = per_instance_rate * 0.022 # ~29 requests at W = 22 ms
The last line is L again, at the level of the instance: about 29 requests in flight at the planned rate, which sizes memory — multiply by memory per in-flight request — and any concurrency limit at the front door. The shared check matters as much as the per-instance one: four instances with 20 connections each need 80 database connections, and a database configured for 50 turns the plan into a connection storm. Adding instances moves the bottleneck to the shared resource; cap total connections there, for example with a pooler, as in Connection Pooling & Keep-Alive.
Verify: the plan lists, for the peak rate, utilisation of every pool, total connections to every shared backend, and in-flight requests per instance.
5. Bound the queue with the same law¶
Beyond the limit, an asyncio service does not refuse work by itself: requests keep arriving, coroutines keep waiting on the pool, and latency grows without bound — 1,103 ms of pool wait at p99 at 1,000 req/s here. Little's law gives the bound to enforce instead: if requests must finish within a deadline D, at most λ_max × D of them can usefully be in flight, and anything beyond that can only time out:
MAX_IN_FLIGHT = int(943 * 0.1) # pool limit x 100 ms latency budget = 94
gate = asyncio.Semaphore(MAX_IN_FLIGHT)
async def handler(request):
if gate.locked():
return web.Response(status=503, headers={"Retry-After": "1"})
async with gate:
return await handle(request)
Requests over the bound are rejected immediately, which keeps latency for admitted ones close to normal and tells clients to back off, rather than letting every request time out together. The same reasoning sizes queue lengths between pipeline stages and limits on outgoing calls, as in bounding user-controlled fan-out. Revisit the numbers whenever holding times change: a query that becomes 20% slower lowers the pool limit by the same 20%, and the knee arrives that much sooner.
Verify: under load above the limit, the service returns 503 quickly and admitted requests stay within the latency budget.
Verification¶
A capacity plan built on Little's law is sound when:
- λ × W matches measured in-flight requests below saturation.
- Each pool's limit is computed from its real holding time, and confirmed by a load test.
- Pools and instances are sized for about 70% utilisation at peak, with shared backends checked in total.
- In-flight requests are bounded at limit × latency budget, with fast rejection beyond it.
Diagnostic Hook: export, per pool, busy slots and acquisition wait time, and compute utilisation as busy ÷ size. Pool wait that rises while utilisation crosses about 85% is the knee approaching; pool wait that rises while utilisation is low means slots are held by something other than the queries you planned for — a slow transaction, a leaked connection, or work done while holding one.
Pitfalls & edge cases¶
- Using the nominal query time. Measured holding time was 10.6 ms, moving the limit from 1,000 to about 943 req/s.
- Planning at 100% utilisation. At 95% busy, pool wait p99 was already 11.5 ms.
- Applying the law to a growing queue. It holds only in steady state.
- Forgetting shared backends. Instances × pool size must fit the database.
Frequently Asked Questions¶
What is Little's law in capacity planning?
The average number of requests in a system equals arrival rate times average time in the system (L = λ × W). For an asyncio service it predicted in-flight requests within 3% of measurement below saturation.
How do I size a database connection pool with Little's law?
Busy connections = request rate × time each request holds a connection. Divide by a target utilisation of about 70%: 800 req/s at 10.6 ms needs about 13 connections.
When will my service saturate?
At pool size ÷ holding time for its scarcest resource. A 10-slot pool at 10.6 ms per request predicted 943 req/s; latency jumped between 900 and 950 req/s, and doubling the pool moved the knee past 1,700.
How many requests should a service allow in flight?
About its saturation rate times the latency budget; more can only wait and time out. Reject the excess quickly with 503.
Related¶
- Load Testing & Benchmarking — up to the topic overview.
- Finding the saturation point of an async service — measuring the knee directly.
- Resilience, Cancellation & Error Handling — the section overview.