Skip to content

Sharing Large Arrays with Shared Memory

Offloading CPU work to a process pool means moving data between processes, and the default mechanism — pickling arguments through a pipe — copies every byte. For a large array that copy is the job. Measured on Python 3.14 with NumPy and eight worker processes, summing eight slices of an 80 MB float array took 1.34 s when the whole array was passed to every worker, 0.157 s when each worker received only its slice, and 0.016 s when the array lived in a multiprocessing.shared_memory block and each worker received just its name and bounds. The in-process sum took 0.009 s, so shared memory brought the pool's overhead close to zero. This guide covers the shared-memory pattern from asyncio and the lifecycle rules that keep it from leaking.

Prerequisites

1. See where the time goes

import asyncio
import numpy as np
from concurrent.futures import ProcessPoolExecutor


def total(arr: np.ndarray, lo: int, hi: int) -> float:
    return float(arr[lo:hi].sum())


async def pickled(pool, data: np.ndarray) -> list[float]:
    loop = asyncio.get_running_loop()
    n = len(data) // 8
    return await asyncio.gather(*(
        loop.run_in_executor(pool, total, data, i * n, (i + 1) * n) for i in range(8)
    ))

Each call pickles the full 80 MB array — the pickled size was 80.0 MB — so eight calls push 640 MB through pipes, for a computation that takes 9 ms in-process. Sending only each worker's slice cuts the copy to 80 MB in total and the time to 0.157 s, which is the first fix to make even before shared memory: never pass more data than the worker needs.

Verify: len(pickle.dumps(args)) for one call; if it is megabytes, the transfer is your bottleneck.

Summing 8 slices of an 80 MB array in 8 processes 4 horizontal bars comparing whole array pickled to each with the others. Summing 8 slices of an 80 MB array in 8 processes whole array pickled to each 1.34 s only its slice pickled 0.157 s SharedMemory, name + bounds 0.016 s in-process, no pool 0.009 s Python 3.14, NumPy, 8 worker processes already started. With shared memory the pool's transfer cost almost disappears; only the work remains.

2. Put the array in a SharedMemory block

Create the block in the parent, copy the data in once, and pass workers the block's name, the shape, the dtype and their bounds — a few dozen bytes:

from multiprocessing import shared_memory


def total_shm(name: str, shape: tuple, dtype: str, lo: int, hi: int) -> float:
    shm = shared_memory.SharedMemory(name=name)            # attach, no copy
    try:
        arr = np.ndarray(shape, dtype=dtype, buffer=shm.buf)
        return float(arr[lo:hi].sum())
    finally:
        shm.close()                                         # detach this process only


async def via_shared_memory(pool, data: np.ndarray) -> list[float]:
    loop = asyncio.get_running_loop()
    shm = shared_memory.SharedMemory(create=True, size=data.nbytes)
    try:
        view = np.ndarray(data.shape, dtype=data.dtype, buffer=shm.buf)
        view[:] = data                                       # one copy, in the parent
        n = len(data) // 8
        return await asyncio.gather(*(
            loop.run_in_executor(pool, total_shm, shm.name, data.shape, data.dtype.str,
                                 i * n, (i + 1) * n)
            for i in range(8)
        ))
    finally:
        del view                                             # release the buffer export
        shm.close()
        shm.unlink()                                         # free the block itself

Measured: 0.016 s, with results identical to the pickled version. Workers attach by name, read through the same physical pages, and detach. Only the creator calls unlink().

Verify: results match the pickled version exactly, and /dev/shm contains no leftover psm_* files after the call.

3. Get the lifecycle right

SharedMemory has two operations that sound alike and are not:

  • close() detaches this process's mapping. Every process that attached must call it.
  • unlink() destroys the block. Exactly one process — the owner — calls it, once, after everyone is done.
from contextlib import contextmanager


@contextmanager
def shared_array(data: np.ndarray):
    shm = shared_memory.SharedMemory(create=True, size=data.nbytes)
    view = np.ndarray(data.shape, dtype=data.dtype, buffer=shm.buf)
    view[:] = data
    try:
        yield shm.name, data.shape, data.dtype.str
    finally:
        del view
        shm.close()
        shm.unlink()

Two traps. A NumPy array created on shm.buf keeps pointing at the mapping, and close() does not stop you: tested on Python 3.12, 3.13 and 3.14, close() succeeded silently with a live view still in scope. The view then refers to unmapped memory, and touching it can crash the process — hence del view before close(), every time. And a process that crashes before unlink() leaves the block in /dev/shm until reboot; on Linux it counts against the container's shared-memory limit (often a small /dev/shm by default in Docker), so leaks eventually make create=True fail. Wrapping creation and unlinking in a context manager or try/finally is not optional.

Verify: run the workload, kill the parent mid-run, and check /dev/shm; then add a startup sweep or unique name prefixes so orphans can be found and removed.

Who creates, attaches, closes and unlinks A sequence of 6 messages between 3 participants. Who creates, attaches, closes and unlinks parent shared block worker create, copy array in send name, shape, dtype, bounds attach by name, read slice close(): detach worker del view, close() unlink(): free the block Every process closes; only the creator unlinks, and only after every worker is done.

4. Write results back without copying

Workers can also write into shared memory — the usual way to return large results without pickling them back. Give each worker a disjoint output region:

def normalise_into(in_name, out_name, shape, dtype, lo, hi) -> None:
    src = shared_memory.SharedMemory(name=in_name)
    dst = shared_memory.SharedMemory(name=out_name)
    try:
        a = np.ndarray(shape, dtype=dtype, buffer=src.buf)
        b = np.ndarray(shape, dtype=dtype, buffer=dst.buf)
        seg = a[lo:hi]
        b[lo:hi] = (seg - seg.mean()) / (seg.std() or 1.0)  # each worker writes only [lo, hi)
        del a, b
    finally:
        src.close()
        dst.close()

Disjoint regions need no locking. Overlapping writes do, and a multiprocessing.Lock around a hot loop erases the speed-up; design the partitioning so overlap never happens. The parent reads the output block after gather() returns — the completion of every future is the synchronisation point.

Verify: the output equals the single-process computation, element for element.

5. Know when shared memory is not worth it

Shared memory removes transfer cost; it adds lifecycle code and makes data layout explicit. Use it when the data is large and read by many tasks, not as a default:

Situation Better choice
inputs under a few hundred KB per call ordinary arguments
large array, each worker needs a slice pass slices — 0.157 s vs 1.34 s
large array read repeatedly by many tasks SharedMemory — 0.016 s
large read-only data loaded once per worker pool initializer loading it from disk
NumPy work that releases the GIL threads, no copying at all

The last row deserves attention: many NumPy operations release the GIL, so a ThreadPoolExecutor can run them in parallel with zero transfer — the subject of using threads for GIL-releasing C extensions. The initializer approach is in initializing process pool workers with expensive state.

Verify: benchmark the simplest option first; reach for shared memory only when the measured transfer cost justifies it.

How should large data reach the workers? A decision on How is the data used with 3 outcomes. How should large data reach the workers? How is the data used? each worker needs a part, once pass slices simplest big win many tasks read the same array SharedMemory near-zero transfer work releases the GIL threads instead no copy at all Copying less beats copying faster; shared memory is the end of that road, not the start.

Verification

Shared-memory offloading is correct when:

  • Calls carry names and bounds, not arrays — pickled argument size is bytes, not megabytes.
  • Every attach is matched by close(), and the owner unlinks exactly once.
  • No psm_* blocks remain in /dev/shm after runs, including failed ones.
  • Writes go to disjoint regions, and results match a single-process run.

Diagnostic Hook: export the number and total size of shared-memory blocks your process has created and not yet unlinked, and alert if it grows. On Linux, also watch /dev/shm usage on the host or container: a slowly filling /dev/shm is the signature of crashed processes leaving blocks behind, and it ends in create=True failures that look unrelated to the original crash.

Pitfalls & edge cases

  • Closing with a live view. close() succeeds silently on 3.12–3.14 and leaves the NumPy view dangling; delete views first.
  • Unlinking from a worker. Destroys the block while other workers may still attach.
  • Small /dev/shm in containers. Docker defaults are small; raise --shm-size for large blocks.
  • Pickling the SharedMemory object itself. Pass the name; workers attach on their own.

Frequently Asked Questions

How do I share a NumPy array between processes without copying?

Create a multiprocessing.shared_memory.SharedMemory block, copy the array into an ndarray backed by its buffer, and pass workers the block name, shape and dtype. Each worker attaches by name and builds an ndarray view on the same memory.

How much faster is shared memory than pickling arrays?

In testing with an 80 MB array and 8 workers, pickling the whole array to each took 1.34 s, pickling only each worker's slice 0.157 s, and shared memory 0.016 s.

What is the difference between SharedMemory.close and unlink?

close() detaches the current process's mapping and must be called by every process that attached. unlink() destroys the block and must be called once, by the owner, after all users are done.

Do I need to delete NumPy views before SharedMemory.close?

Yes. On Python 3.12 to 3.14, close() succeeded in testing even with a NumPy view still alive, leaving the view pointing at unmapped memory that can crash the process when touched. Delete views first.