Validating Large Request Bodies Off the Event Loop¶
FastAPI validates a typed request body on the event loop, and for a large body that validation is long enough to stall every other request in the process. Measured on Python 3.14 with FastAPI 0.142, Pydantic 2.13 and uvicorn 0.54, posting a 3.2 MB JSON order with 40,000 line items from 4 concurrent clients while a probe timed a trivial /ping route: validating the body took 132 ms per request, the typed endpoint served 7.3 requests per second, and /ping had a median of 145 ms and a p99 of 307 ms. Moving validation to asyncio.to_thread barely helped — p50 132 ms, p99 156 ms — because Pydantic's validator holds the GIL while it runs. A four-process pool brought /ping down to 0.6 ms median and 2.6 ms p99 and raised throughput to 19.3 requests per second — as long as the worker returned a small result. Returning the validated model instead cost 251 ms to pickle and 154 ms to unpickle, cut throughput to 5.8 per second, and pushed /ping p99 back to 51 ms. This guide measures each option and builds the version that works.
Prerequisites¶
- FastAPI with Pydantic v2 and uvicorn.
- Body size limits, from limiting request body size in ASGI apps.
- The topic overview, ASGI Servers & Frameworks.
1. Measure validation cost outside the server¶
Time the validator on a realistic large body before touching the app:
class Item(BaseModel):
id: int
sku: str
qty: int
price: float
tags: list[str]
class Order(BaseModel):
customer: str
items: list[Item]
OrderTA = TypeAdapter(Order)
t = time.perf_counter()
order = OrderTA.validate_json(raw) # raw: 3.2 MB, 40,000 items
print((time.perf_counter() - t) * 1000)
Measured: 132 ms with validate_json, and 122 ms with json.loads followed by Order.model_validate — the two-step path was not slower here, so measure both on your own model. Either way, 132 ms is 13 times the 10 ms that most services can tolerate as a single blocking call, and a request handler doing nothing else is blocked for all of it.
Verify: validation time for the largest body the endpoint accepts is known, and compared with the latency budget of the other routes in the process.
2. Measure the typed endpoint's effect on other routes¶
The idiomatic FastAPI endpoint declares the model as a parameter, and the framework parses and validates the body before the handler runs:
@app.post("/orders")
async def create_order(order: Order):
return {"n": len(order.items)}
Measured with 4 clients posting the 3.2 MB body continuously: 7.3 requests per second, each taking 582 ms at the median as they queued for the loop, and /ping — a route that returns a constant — at a 145 ms median and 307 ms p99. The server process was at 96% of one core. Reading the raw body and calling the validator in the handler behaved the same way: 6.5 requests per second and a /ping p99 of 422 ms. Validation runs where the framework's code runs, on the event loop, and every other request on that worker waits for it.
Verify: run a large-body load test with a probe on a trivial route; a probe p99 in the hundreds of milliseconds means validation is on the loop.
3. Check whether a thread helps¶
asyncio.to_thread helps only when the work releases the GIL. Test it rather than assuming:
@app.post("/orders")
async def create_order(request: Request):
raw = await request.body()
n = await asyncio.to_thread(validate, raw) # validate: OrderTA.validate_json(raw)
return {"n": n}
Measured: 7.5 requests per second, and /ping at 132 ms median and 156 ms p99. The p99 improved from 307 ms, because the interpreter forces a GIL handoff every 5 ms, but the median barely moved: while the validation thread ran, the event loop thread waited for the GIL. Pydantic's Rust core does its work while holding the GIL, so threads give it neither parallelism nor responsiveness. Contrast zlib compression, which releases the GIL and gains a great deal from threads, as measured in compressing responses without blocking the loop.
Verify: the probe's median, not just its p99, drops after moving work to a thread; if it does not, the work holds the GIL.
4. Validate in a process pool and return little¶
A process has its own GIL. Create the pool in the app's lifespan, send the raw bytes, and return only what the handler needs:
def validate_and_summarise(raw: bytes) -> dict:
order = OrderTA.validate_json(raw) # runs in a worker process
return {"customer": order.customer, "n": len(order.items),
"total": sum(i.qty * i.price for i in order.items)}
@asynccontextmanager
async def lifespan(app):
app.state.pool = ProcessPoolExecutor(max_workers=4)
yield
app.state.pool.shutdown(wait=True, cancel_futures=True)
app = FastAPI(lifespan=lifespan)
@app.post("/orders")
async def create_order(request: Request):
raw = await request.body()
loop = asyncio.get_running_loop()
return await loop.run_in_executor(request.app.state.pool, validate_and_summarise, raw)
Measured with 4 clients: 19.3 requests per second at a 209 ms median, and /ping at 0.6 ms median and 2.6 ms p99. With 16 clients, throughput stayed at 19.2 — four processes at about 200 ms each — and the probe's p99 at 4.9 ms. What the worker returns matters as much as where it runs. Returning the validated Order instead meant pickling 40,000 model instances — 251 ms in the worker — and unpickling them on the event loop in 154 ms. That variant fell to 5.8 requests per second with a /ping p99 of 51 ms: the blocking had moved from validation to deserialisation. Do the work that needs the full model in the worker too, such as computing totals or writing rows, and send back a summary.
Verify: the result returned from the pool is small, and its unpickling time on the loop is measured in microseconds.
5. Shut the pool down and bound the input¶
Two operational details complete the pattern. First, shut the pool down in the lifespan, as above: with Python 3.14's default forkserver start method on Linux, the pool's processes are not children of the server process. Measured: stopping uvicorn with SIGTERM when the pool was created at import time left the forkserver, the resource tracker and the four workers running — six processes per server instance. With pool.shutdown() in the lifespan, none remained.
app.add_middleware(MaxBodySizeMiddleware, max_bytes=4 * 1024 * 1024) # reject before reading
Second, cap the body size before validation, since the cost grows with it: a 6.5 MB body with 80,000 items took 247 ms against 120 ms for 3.2 MB in a repeat measurement, and an unbounded endpoint lets one client occupy a worker indefinitely. A size limit rejects oversize bodies with 413 before they are read, as built in limiting request body size in ASGI apps. Small bodies — a few KiB — should keep using typed parameters; a pool round trip costs more than validating them inline.
Verify: after stopping the server, no pool processes remain; and a body above the limit is rejected with 413 without being validated.
Verification¶
Large-body validation is handled when:
- Validation time per body is measured for the largest accepted size.
- A probe route stays responsive — p99 of a few milliseconds — during a large-body load test.
- Workers return summaries, not validated models.
- The pool is shut down in the lifespan, and bodies above a size limit are rejected first.
Diagnostic Hook: when all routes in a FastAPI worker slow down together during uploads, compare the probe's median with the upload endpoint's validation time. A probe median close to it — 145 ms against 132 ms here — means validation is running on the event loop.
Pitfalls & edge cases¶
- Typed parameters for multi-megabyte bodies. Measured:
/pingp99 307 ms. to_threadfor Pydantic validation. Measured: median probe latency stayed at 132 ms.- Returning the model from the pool. Measured: 405 ms of pickling round trip and 5.8 requests/s.
- Pools created at import time. Six processes per server outlived a SIGTERM.
Frequently Asked Questions¶
Does FastAPI validate request bodies on the event loop?
Yes. A typed body parameter is parsed and validated before the handler, on the loop. A 3.2 MB body took 132 ms and pushed a trivial route's p99 to 307 ms.
Does Pydantic release the GIL during validation?
No, in these measurements: validating in asyncio.to_thread left a trivial route's median latency at 132 ms, the same as the validation time.
How do I validate large JSON bodies without blocking FastAPI?
Read the raw body, validate it in a ProcessPoolExecutor created in the lifespan, and return a small summary. The trivial route's p99 fell to 2.6 ms.
Why is returning a Pydantic model from a process pool slow?
It is pickled in the worker and unpickled on the loop: 251 ms and 154 ms for 40,000 items. Do the work in the worker and return a small result.
Related¶
- ASGI Servers & Frameworks — up to the topic overview.
- Serving static files from ASGI — another route type that loads the server process.
- Network I/O & Protocol Handling — the section overview.