Skip to content

Validating Large Request Bodies Off the Event Loop

FastAPI validates a typed request body on the event loop, and for a large body that validation is long enough to stall every other request in the process. Measured on Python 3.14 with FastAPI 0.142, Pydantic 2.13 and uvicorn 0.54, posting a 3.2 MB JSON order with 40,000 line items from 4 concurrent clients while a probe timed a trivial /ping route: validating the body took 132 ms per request, the typed endpoint served 7.3 requests per second, and /ping had a median of 145 ms and a p99 of 307 ms. Moving validation to asyncio.to_thread barely helped — p50 132 ms, p99 156 ms — because Pydantic's validator holds the GIL while it runs. A four-process pool brought /ping down to 0.6 ms median and 2.6 ms p99 and raised throughput to 19.3 requests per second — as long as the worker returned a small result. Returning the validated model instead cost 251 ms to pickle and 154 ms to unpickle, cut throughput to 5.8 per second, and pushed /ping p99 back to 51 ms. This guide measures each option and builds the version that works.

Prerequisites

1. Measure validation cost outside the server

Time the validator on a realistic large body before touching the app:

class Item(BaseModel):
    id: int
    sku: str
    qty: int
    price: float
    tags: list[str]

class Order(BaseModel):
    customer: str
    items: list[Item]

OrderTA = TypeAdapter(Order)

t = time.perf_counter()
order = OrderTA.validate_json(raw)          # raw: 3.2 MB, 40,000 items
print((time.perf_counter() - t) * 1000)

Measured: 132 ms with validate_json, and 122 ms with json.loads followed by Order.model_validate — the two-step path was not slower here, so measure both on your own model. Either way, 132 ms is 13 times the 10 ms that most services can tolerate as a single blocking call, and a request handler doing nothing else is blocked for all of it.

Verify: validation time for the largest body the endpoint accepts is known, and compared with the latency budget of the other routes in the process.

2. Measure the typed endpoint's effect on other routes

The idiomatic FastAPI endpoint declares the model as a parameter, and the framework parses and validates the body before the handler runs:

@app.post("/orders")
async def create_order(order: Order):
    return {"n": len(order.items)}

Measured with 4 clients posting the 3.2 MB body continuously: 7.3 requests per second, each taking 582 ms at the median as they queued for the loop, and /ping — a route that returns a constant — at a 145 ms median and 307 ms p99. The server process was at 96% of one core. Reading the raw body and calling the validator in the handler behaved the same way: 6.5 requests per second and a /ping p99 of 422 ms. Validation runs where the framework's code runs, on the event loop, and every other request on that worker waits for it.

Verify: run a large-body load test with a probe on a trivial route; a probe p99 in the hundreds of milliseconds means validation is on the loop.

3.2 MB order bodies, 4 concurrent clients A grid of 5 rows by 4 columns. 3.2 MB order bodies, 4 concurrent clients approach orders/s /ping p50 /ping p99 typed parameter (FastAPI default) 7.3 145 ms 307 ms raw body, validate inline 6.5 182 ms 422 ms raw body, asyncio.to_thread 7.5 132 ms 156 ms process pool, returns a count 19.3 0.6 ms 2.6 ms process pool, returns the model 5.8 18.0 ms 51.0 ms FastAPI 0.142.2, Pydantic 2.13.5, uvicorn 0.54, Python 3.14.

3. Check whether a thread helps

asyncio.to_thread helps only when the work releases the GIL. Test it rather than assuming:

@app.post("/orders")
async def create_order(request: Request):
    raw = await request.body()
    n = await asyncio.to_thread(validate, raw)       # validate: OrderTA.validate_json(raw)
    return {"n": n}

Measured: 7.5 requests per second, and /ping at 132 ms median and 156 ms p99. The p99 improved from 307 ms, because the interpreter forces a GIL handoff every 5 ms, but the median barely moved: while the validation thread ran, the event loop thread waited for the GIL. Pydantic's Rust core does its work while holding the GIL, so threads give it neither parallelism nor responsiveness. Contrast zlib compression, which releases the GIL and gains a great deal from threads, as measured in compressing responses without blocking the loop.

Verify: the probe's median, not just its p99, drops after moving work to a thread; if it does not, the work holds the GIL.

4. Validate in a process pool and return little

A process has its own GIL. Create the pool in the app's lifespan, send the raw bytes, and return only what the handler needs:

def validate_and_summarise(raw: bytes) -> dict:
    order = OrderTA.validate_json(raw)              # runs in a worker process
    return {"customer": order.customer, "n": len(order.items),
            "total": sum(i.qty * i.price for i in order.items)}

@asynccontextmanager
async def lifespan(app):
    app.state.pool = ProcessPoolExecutor(max_workers=4)
    yield
    app.state.pool.shutdown(wait=True, cancel_futures=True)

app = FastAPI(lifespan=lifespan)

@app.post("/orders")
async def create_order(request: Request):
    raw = await request.body()
    loop = asyncio.get_running_loop()
    return await loop.run_in_executor(request.app.state.pool, validate_and_summarise, raw)

Measured with 4 clients: 19.3 requests per second at a 209 ms median, and /ping at 0.6 ms median and 2.6 ms p99. With 16 clients, throughput stayed at 19.2 — four processes at about 200 ms each — and the probe's p99 at 4.9 ms. What the worker returns matters as much as where it runs. Returning the validated Order instead meant pickling 40,000 model instances — 251 ms in the worker — and unpickling them on the event loop in 154 ms. That variant fell to 5.8 requests per second with a /ping p99 of 51 ms: the blocking had moved from validation to deserialisation. Do the work that needs the full model in the worker too, such as computing totals or writing rows, and send back a summary.

Verify: the result returned from the pool is small, and its unpickling time on the loop is measured in microseconds.

Validating a body in a process pool A sequence of 6 messages between 3 participants. Validating a body in a process pool client event loop pool worker POST /orders, 3.2 MB await request.body(), no parsing run_in_executor(raw bytes) validate_json: 132 ms small dict, not the model 200 + summary The loop serves other routes while the worker validates.

5. Shut the pool down and bound the input

Two operational details complete the pattern. First, shut the pool down in the lifespan, as above: with Python 3.14's default forkserver start method on Linux, the pool's processes are not children of the server process. Measured: stopping uvicorn with SIGTERM when the pool was created at import time left the forkserver, the resource tracker and the four workers running — six processes per server instance. With pool.shutdown() in the lifespan, none remained.

app.add_middleware(MaxBodySizeMiddleware, max_bytes=4 * 1024 * 1024)   # reject before reading

Second, cap the body size before validation, since the cost grows with it: a 6.5 MB body with 80,000 items took 247 ms against 120 ms for 3.2 MB in a repeat measurement, and an unbounded endpoint lets one client occupy a worker indefinitely. A size limit rejects oversize bodies with 413 before they are read, as built in limiting request body size in ASGI apps. Small bodies — a few KiB — should keep using typed parameters; a pool round trip costs more than validating them inline.

Verify: after stopping the server, no pool processes remain; and a body above the limit is rejected with 413 without being validated.

/ping p99 while 3.2 MB orders are validated 5 horizontal bars comparing validate inline (raw body) with the others. /ping p99 while 3.2 MB orders are validated validate inline (raw body) 422 ms typed parameter 307 ms asyncio.to_thread 156 ms process pool, returns the model 51 ms process pool, returns a summary 2.6 ms Only a process pool with a small result kept the loop responsive.

Verification

Large-body validation is handled when:

  • Validation time per body is measured for the largest accepted size.
  • A probe route stays responsive — p99 of a few milliseconds — during a large-body load test.
  • Workers return summaries, not validated models.
  • The pool is shut down in the lifespan, and bodies above a size limit are rejected first.

Diagnostic Hook: when all routes in a FastAPI worker slow down together during uploads, compare the probe's median with the upload endpoint's validation time. A probe median close to it — 145 ms against 132 ms here — means validation is running on the event loop.

Pitfalls & edge cases

  • Typed parameters for multi-megabyte bodies. Measured: /ping p99 307 ms.
  • to_thread for Pydantic validation. Measured: median probe latency stayed at 132 ms.
  • Returning the model from the pool. Measured: 405 ms of pickling round trip and 5.8 requests/s.
  • Pools created at import time. Six processes per server outlived a SIGTERM.

Frequently Asked Questions

Does FastAPI validate request bodies on the event loop?

Yes. A typed body parameter is parsed and validated before the handler, on the loop. A 3.2 MB body took 132 ms and pushed a trivial route's p99 to 307 ms.

Does Pydantic release the GIL during validation?

No, in these measurements: validating in asyncio.to_thread left a trivial route's median latency at 132 ms, the same as the validation time.

How do I validate large JSON bodies without blocking FastAPI?

Read the raw body, validate it in a ProcessPoolExecutor created in the lifespan, and return a small summary. The trivial route's p99 fell to 2.6 ms.

Why is returning a Pydantic model from a process pool slow?

It is pickled in the worker and unpickled on the loop: 251 ms and 154 ms for 40,000 items. Do the work in the worker and return a small result.