Sharing Large Arrays with Shared Memory¶
Offloading CPU work to a process pool means moving data between processes, and the default mechanism — pickling arguments through a pipe — copies every byte. For a large array that copy is the job. Measured on Python 3.14 with NumPy and eight worker processes, summing eight slices of an 80 MB float array took 1.34 s when the whole array was passed to every worker, 0.157 s when each worker received only its slice, and 0.016 s when the array lived in a multiprocessing.shared_memory block and each worker received just its name and bounds. The in-process sum took 0.009 s, so shared memory brought the pool's overhead close to zero. This guide covers the shared-memory pattern from asyncio and the lifecycle rules that keep it from leaking.
Prerequisites¶
- Python 3.11+; examples use
pip install numpy, butSharedMemoryworks with any buffer, includingbytearrayandmemoryview. - Process pools from asyncio, from offloading CPU work with run_in_executor.
- Pickle costs, from reducing pickle overhead in ProcessPoolExecutor payloads.
1. See where the time goes¶
import asyncio
import numpy as np
from concurrent.futures import ProcessPoolExecutor
def total(arr: np.ndarray, lo: int, hi: int) -> float:
return float(arr[lo:hi].sum())
async def pickled(pool, data: np.ndarray) -> list[float]:
loop = asyncio.get_running_loop()
n = len(data) // 8
return await asyncio.gather(*(
loop.run_in_executor(pool, total, data, i * n, (i + 1) * n) for i in range(8)
))
Each call pickles the full 80 MB array — the pickled size was 80.0 MB — so eight calls push 640 MB through pipes, for a computation that takes 9 ms in-process. Sending only each worker's slice cuts the copy to 80 MB in total and the time to 0.157 s, which is the first fix to make even before shared memory: never pass more data than the worker needs.
Verify: len(pickle.dumps(args)) for one call; if it is megabytes, the transfer is your bottleneck.
2. Put the array in a SharedMemory block¶
Create the block in the parent, copy the data in once, and pass workers the block's name, the shape, the dtype and their bounds — a few dozen bytes:
from multiprocessing import shared_memory
def total_shm(name: str, shape: tuple, dtype: str, lo: int, hi: int) -> float:
shm = shared_memory.SharedMemory(name=name) # attach, no copy
try:
arr = np.ndarray(shape, dtype=dtype, buffer=shm.buf)
return float(arr[lo:hi].sum())
finally:
shm.close() # detach this process only
async def via_shared_memory(pool, data: np.ndarray) -> list[float]:
loop = asyncio.get_running_loop()
shm = shared_memory.SharedMemory(create=True, size=data.nbytes)
try:
view = np.ndarray(data.shape, dtype=data.dtype, buffer=shm.buf)
view[:] = data # one copy, in the parent
n = len(data) // 8
return await asyncio.gather(*(
loop.run_in_executor(pool, total_shm, shm.name, data.shape, data.dtype.str,
i * n, (i + 1) * n)
for i in range(8)
))
finally:
del view # release the buffer export
shm.close()
shm.unlink() # free the block itself
Measured: 0.016 s, with results identical to the pickled version. Workers attach by name, read through the same physical pages, and detach. Only the creator calls unlink().
Verify: results match the pickled version exactly, and /dev/shm contains no leftover psm_* files after the call.
3. Get the lifecycle right¶
SharedMemory has two operations that sound alike and are not:
close()detaches this process's mapping. Every process that attached must call it.unlink()destroys the block. Exactly one process — the owner — calls it, once, after everyone is done.
from contextlib import contextmanager
@contextmanager
def shared_array(data: np.ndarray):
shm = shared_memory.SharedMemory(create=True, size=data.nbytes)
view = np.ndarray(data.shape, dtype=data.dtype, buffer=shm.buf)
view[:] = data
try:
yield shm.name, data.shape, data.dtype.str
finally:
del view
shm.close()
shm.unlink()
Two traps. A NumPy array created on shm.buf keeps pointing at the mapping, and close() does not stop you: tested on Python 3.12, 3.13 and 3.14, close() succeeded silently with a live view still in scope. The view then refers to unmapped memory, and touching it can crash the process — hence del view before close(), every time. And a process that crashes before unlink() leaves the block in /dev/shm until reboot; on Linux it counts against the container's shared-memory limit (often a small /dev/shm by default in Docker), so leaks eventually make create=True fail. Wrapping creation and unlinking in a context manager or try/finally is not optional.
Verify: run the workload, kill the parent mid-run, and check /dev/shm; then add a startup sweep or unique name prefixes so orphans can be found and removed.
4. Write results back without copying¶
Workers can also write into shared memory — the usual way to return large results without pickling them back. Give each worker a disjoint output region:
def normalise_into(in_name, out_name, shape, dtype, lo, hi) -> None:
src = shared_memory.SharedMemory(name=in_name)
dst = shared_memory.SharedMemory(name=out_name)
try:
a = np.ndarray(shape, dtype=dtype, buffer=src.buf)
b = np.ndarray(shape, dtype=dtype, buffer=dst.buf)
seg = a[lo:hi]
b[lo:hi] = (seg - seg.mean()) / (seg.std() or 1.0) # each worker writes only [lo, hi)
del a, b
finally:
src.close()
dst.close()
Disjoint regions need no locking. Overlapping writes do, and a multiprocessing.Lock around a hot loop erases the speed-up; design the partitioning so overlap never happens. The parent reads the output block after gather() returns — the completion of every future is the synchronisation point.
Verify: the output equals the single-process computation, element for element.
5. Know when shared memory is not worth it¶
Shared memory removes transfer cost; it adds lifecycle code and makes data layout explicit. Use it when the data is large and read by many tasks, not as a default:
| Situation | Better choice |
|---|---|
| inputs under a few hundred KB per call | ordinary arguments |
| large array, each worker needs a slice | pass slices — 0.157 s vs 1.34 s |
| large array read repeatedly by many tasks | SharedMemory — 0.016 s |
| large read-only data loaded once per worker | pool initializer loading it from disk |
| NumPy work that releases the GIL | threads, no copying at all |
The last row deserves attention: many NumPy operations release the GIL, so a ThreadPoolExecutor can run them in parallel with zero transfer — the subject of using threads for GIL-releasing C extensions. The initializer approach is in initializing process pool workers with expensive state.
Verify: benchmark the simplest option first; reach for shared memory only when the measured transfer cost justifies it.
Verification¶
Shared-memory offloading is correct when:
- Calls carry names and bounds, not arrays — pickled argument size is bytes, not megabytes.
- Every attach is matched by
close(), and the owner unlinks exactly once. - No
psm_*blocks remain in/dev/shmafter runs, including failed ones. - Writes go to disjoint regions, and results match a single-process run.
Diagnostic Hook: export the number and total size of shared-memory blocks your process has created and not yet unlinked, and alert if it grows. On Linux, also watch /dev/shm usage on the host or container: a slowly filling /dev/shm is the signature of crashed processes leaving blocks behind, and it ends in create=True failures that look unrelated to the original crash.
Pitfalls & edge cases¶
- Closing with a live view.
close()succeeds silently on 3.12–3.14 and leaves the NumPy view dangling; delete views first. - Unlinking from a worker. Destroys the block while other workers may still attach.
- Small
/dev/shmin containers. Docker defaults are small; raise--shm-sizefor large blocks. - Pickling the
SharedMemoryobject itself. Pass the name; workers attach on their own.
Frequently Asked Questions¶
How do I share a NumPy array between processes without copying?
Create a multiprocessing.shared_memory.SharedMemory block, copy the array into an ndarray backed by its buffer, and pass workers the block name, shape and dtype. Each worker attaches by name and builds an ndarray view on the same memory.
How much faster is shared memory than pickling arrays?
In testing with an 80 MB array and 8 workers, pickling the whole array to each took 1.34 s, pickling only each worker's slice 0.157 s, and shared memory 0.016 s.
What is the difference between SharedMemory.close and unlink?
close() detaches the current process's mapping and must be called by every process that attached. unlink() destroys the block and must be called once, by the owner, after all users are done.
Do I need to delete NumPy views before SharedMemory.close?
Yes. On Python 3.12 to 3.14, close() succeeded in testing even with a NumPy view still alive, leaving the view pointing at unmapped memory that can crash the process when touched. Delete views first.
Related¶
- CPU-Bound Task Offloading — up to the topic overview.
- Choosing chunksize for ProcessPoolExecutor work — the per-call overhead side of the same problem.
- Concurrent Execution & Worker Patterns — the section overview.