Load Testing Async Services with Locust¶
Locust is the most common way to load test a Python service, and its results are easy to misread: a user count is not a request rate, and a single Locust process can saturate before the service does, at which point the latency it reports is its own. Measured with Locust 2.46.6 against a FastAPI service on four uvicorn workers whose endpoint awaits a 20 ms downstream call, with Locust pinned to a single core: 50 users with a 0.5–1 s wait between requests generated only 66 req/s. Aiming for 3,000 req/s with constant_throughput, HttpUser reached 1,860–1,933 req/s with Locust at 99% CPU, reporting a median of 41–42 ms for an endpoint that takes about 21 ms. FastHttpUser reached 2,985–2,994 req/s at 63% CPU with a median of 21–22 ms, and HttpUser spread over four processes reached 2,941 req/s at a median of 23 ms. An SLO check of p95 ≤ 50 ms failed with the saturated generator (57 ms) and passed with FastHttpUser (31 ms) against the same service. This guide sets up Locust so its numbers describe the service.
Prerequisites¶
- Locust 2.x in its own environment, on a machine or cores separate from the service.
- Open versus closed load, from avoiding coordinated omission in latency benchmarks.
- The topic overview, Load Testing & Benchmarking.
1. Specify load as a rate, not a user count¶
A Locust user runs its tasks in a loop with a wait_time between them, so the request rate depends on how long each request takes. With a think time, it is mostly set by the think time:
from locust import HttpUser, task, between
class Browsing(HttpUser):
wait_time = between(0.5, 1.0)
@task
def get_item(self):
self.client.get("/item/42", name="/item/[id]")
Measured: 50 users produced 65.9 req/s — about one request per user every 0.76 s, the mean think time plus the 21 ms response. That is the right model for "50 people browsing", and the wrong one for "the service must handle 1,000 req/s". For a target rate, use constant_throughput, which paces each user to a number of task runs per second:
from locust import FastHttpUser, task, constant_throughput
class Api(FastHttpUser):
wait_time = constant_throughput(10) # 10 req/s per user; 100 users -> 1,000 req/s
@task
def get_item(self):
self.client.get("/item/42", name="/item/[id]")
With 100 users at 10 each, both user classes delivered about 1,000 req/s. constant_throughput is still closed-loop: a user whose request is slow sends its next one late, so a stall reduces the load instead of queueing it, and the latency during the stall is under-reported. For stall-sensitive measurements, pair Locust with an open-loop generator as in building an open-loop load generator in asyncio.
Verify: the achieved request rate in Locust's summary matches the rate you meant to apply.
2. Check that Locust is not the bottleneck¶
Locust measures response time on its side, including any time a request waits for Locust itself. When the Locust process runs out of CPU, requests are sent late and responses are read late, and both show up as service latency:
/usr/bin/time -f "locust CPU %P" \
locust -f locustfile.py --headless -u 100 -r 100 -t 20s -H http://service:8811 --csv run
Measured with a 3,000 req/s target on one core: HttpUser, built on requests, reached 1,860–1,933 req/s with the Locust process at 99% CPU, and reported p50 41–42 ms and p99 73–77 ms. The service had not changed: the same target with FastHttpUser reported 21–22 ms. Treat Locust's CPU like a gauge on the test itself — above about 80%, its latency numbers include its own queueing. Locust also prints a warning when its CPU is high; read the log, not just the summary.
Verify: the load generator's CPU stays well below 100% at the highest rate you report.
3. Use FastHttpUser, and more processes when needed¶
FastHttpUser uses a lighter HTTP client than requests and does more per core. The task code is the same:
from locust import FastHttpUser
class Api(FastHttpUser):
wait_time = constant_throughput(30)
...
Measured with 100 users targeting 3,000 req/s: 2,985–2,994 req/s achieved at 63% CPU, with a median of 21–22 ms, which matches the service. Where a test needs HttpUser — for a requests-specific auth plugin, say — spread it over several processes with --processes, or run a master with several workers on separate machines:
locust -f locustfile.py --headless -u 100 -r 100 -t 20s --processes 4 -H http://service:8811
Four processes of HttpUser reached 2,941 req/s with a median of 23 ms, using 163% CPU across the four. Both fixes work; the important step is noticing that one was needed. Note that Locust users are gevent greenlets, not asyncio tasks — the locustfile is ordinary blocking-style code, and asyncio libraries cannot be awaited inside it.
Verify: at the target rate, the reported median is close to the service's own measured latency.
4. Turn the run into a pass or fail result¶
A load test in CI needs a verdict. Add a listener that checks the totals when Locust stops and sets the process exit code:
from locust import events
@events.quitting.add_listener
def enforce_slo(environment, **kwargs):
stats = environment.stats.total
p95 = stats.get_response_time_percentile(0.95)
if stats.fail_ratio > 0.01:
environment.process_exit_code = 1
elif p95 > 50:
environment.process_exit_code = 1
else:
environment.process_exit_code = 0
Measured against the same service at 3,000 req/s targeted: with one HttpUser process, the check reported "p95 57 ms > 50 ms" and exited 1; with FastHttpUser, "p95 31 ms" and exited 0. The first result is a false failure: the service was identical, and only the generator was saturated. That is why the generator's CPU belongs in the verdict too — fail the run as invalid, not as a service regression, when the load generator was the bottleneck. Check percentiles from the run's merged histogram, never by averaging per-worker percentiles, for the reasons in measuring latency percentiles without averaging them.
Verify: the run exits non-zero when the SLO is missed, and is marked invalid when the generator's CPU was saturated.
5. Ramp to find limits, then hold to find leaks¶
A single fixed rate answers one question. Locust's load shapes run a sequence of stages, so one test can ramp to find the saturation point and then hold steady to look for slow degradation:
from locust import LoadTestShape
class RampThenHold(LoadTestShape):
stages = [(60, 50), (120, 100), (180, 200), (900, 150)] # (end time s, users)
def tick(self):
t = self.get_run_time()
for end, users in self.stages:
if t < end:
return users, users
return None
With constant_throughput(10), each stage's user count is a request rate — here 500, 1,000 and 2,000 req/s, then 1,500 for twelve minutes. Watch for the stage where latency bends upward, as described in finding the saturation point of an async service, and during the hold watch the service's memory and event-loop lag, which Locust does not see. Record the Locust version, user class, process count and generator CPU with each result, so a later comparison compares services and not test setups.
Verify: the shape's stages produce the intended rates, and service-side metrics are recorded alongside Locust's.
Verification¶
A Locust test measures the service when:
- Load is defined as a rate with
constant_throughput, and the achieved rate matches it. - The generator has CPU headroom —
FastHttpUser, more processes, or more machines. - The reported median matches the service's own latency at low load.
- An SLO listener sets the exit code, and runs with a saturated generator are rejected as invalid.
Diagnostic Hook: run every load test once at a low rate first and record the median. If the median at full rate rises while the service's own request-duration metric does not, the extra latency is in the load generator or the network between them — not something to fix in the service.
Pitfalls & edge cases¶
- Users as a load target. Measured: 50 users produced 66 req/s.
- A saturated generator. Measured: 99% CPU and a median twice the service's.
- Trusting an SLO verdict blindly. It failed because of the generator, not the service.
- Expecting open-loop behaviour.
constant_throughputslows down when the service does.
Frequently Asked Questions¶
How many Locust users do I need for 1,000 requests per second?
Users do not map to a rate by themselves. Use wait_time = constant_throughput(n) so each user makes n requests per second: 100 users at 10 produced about 1,000 req/s in testing.
Why does Locust report higher latency than my service?
Often because the Locust process is out of CPU. One HttpUser process at 99% CPU reported a 42 ms median for a 21 ms endpoint; FastHttpUser reported 21 to 22 ms.
Should I use HttpUser or FastHttpUser?
FastHttpUser for throughput: it reached about 2,990 req/s at 63% CPU on one core where HttpUser stopped near 1,900. Use HttpUser with several processes if you need requests-specific features.
How do I make a Locust test fail in CI?
Add an events.quitting listener that checks environment.stats.total (fail ratio, percentiles) and sets environment.process_exit_code.
Related¶
- Load Testing & Benchmarking — up to the topic overview.
- Load testing WebSocket servers — when the unit of load is a connection.
- Resilience, Cancellation & Error Handling — the section overview.