Skip to content

Raising File Descriptor Limits for Many Connections

Every socket, file and pipe a process holds is a file descriptor, and Linux limits how many one process may have open. The soft limit is often 1,024 — far below what an asyncio server can otherwise handle — and what happens when a server reaches it is worse than a refused connection. Measured on Python 3.14 with an asyncio.start_server echo server whose soft limit was set to 1,024, and 2,000 clients that each connected, exchanged a line and held the connection for 8 s: the server accepted about 1,015 connections, then asyncio logged "socket.accept() out of system resource" — 172,072 times in about 10 seconds, 64 MB of log, with a listen backlog of 4,096. A background task that opened a file every half second failed 12 times with Too many open files. Clients beyond the limit waited in the kernel's queue until others disconnected, pushing p99 connect-and-reply time to 8.8 s; with a backlog of 100, 882 of the 2,000 timed out instead. With the soft limit raised to the hard limit of 524,288, all 2,000 were served with p99 216 ms and no errors. Capping connections at 900 in the handler, with the limit left at 1,024, rejected 1,100 promptly and kept accept errors and file failures at zero. This guide covers raising the limit and staying below it.

Prerequisites

1. Check the limits your process actually has

Each process has a soft limit, which it can raise up to a hard limit, and the limits are inherited from whatever started it — a shell, systemd, a container runtime. Check from inside the service, not from your terminal:

import resource, os

soft, hard = resource.getrlimit(resource.RLIMIT_NOFILE)
in_use = len(os.listdir("/proc/self/fd"))
print(f"file descriptors: {in_use} in use, soft limit {soft}, hard limit {hard}")

In the test environment, the soft and hard limits were both 524,288; many distributions and container defaults still start services with a soft limit of 1,024. The number that matters is the soft limit of the running process — cat /proc/<pid>/limits shows it for any process. Count what the service needs: one descriptor per client connection, plus its database pool, outgoing HTTP connections, log files, and a handful for the loop itself — the idle test server used 8.

Verify: your service logs its descriptor limits at start-up, and the soft limit exceeds peak connections plus everything else it opens.

2. See what happens at the limit

At the limit, accept() fails with EMFILE. asyncio's accept loop logs the failure through the loop's exception handler, stops watching the listening socket and schedules a retry one second later. Measured, the effect was far noisier than that suggests:

asyncio ERROR socket.accept() out of system resource
socket: <asyncio.TransportSocket fd=6, family=2, type=1, proto=6, laddr=('127.0.0.1', 8820)>
OSError: [Errno 24] Too many open files

That message, with its traceback, was logged 172,072 times in about ten seconds. The accept loop in CPython 3.14 tries accept() up to backlog + 1 times per readiness event and, on EMFILE, logs and schedules a retry without leaving the loop — so with a backlog of 4,096, every retry produced thousands of log records. Meanwhile everything else that needed a descriptor failed: the background task's open() raised Too many open files 12 times, and a statistics task that listed /proc/self/fd died. A database reconnect, a log file rotation or a DNS lookup would have failed the same way. The clients themselves were not refused: with a large backlog they waited in the kernel queue until earlier clients disconnected 8 s later; with a backlog of 100, the queue overflowed and 882 of 2,000 clients timed out.

Verify: a load test at more connections than the soft limit shows whether your service degrades this way.

2,000 clients holding connections for 8 s A grid of 4 rows by 5 columns. 2,000 clients holding connections for 8 s server setup clients served p99 accept errors logged open() failures soft limit 1,024, backlog 4,096 2,000 (late) 8,784 ms 172,072 (64 MB) 12 soft limit 1,024, backlog 100 1,118; 882 timed out 9,021 ms 9,392 - soft limit raised to 524,288 2,000 216 ms 0 0 soft limit 1,024, cap 900 in handler 900; 1,100 closed 244 ms 0 0 Python 3.14 asyncio.start_server; client-side timeout 10 s.

3. Raise the soft limit at start-up

A process may raise its own soft limit up to its hard limit without privileges. Do it first thing, before opening anything:

import resource

def raise_nofile_limit(target: int | None = None) -> tuple[int, int]:
    soft, hard = resource.getrlimit(resource.RLIMIT_NOFILE)
    new_soft = hard if target is None else min(target, hard)
    if new_soft > soft:
        resource.setrlimit(resource.RLIMIT_NOFILE, (new_soft, hard))
    return resource.getrlimit(resource.RLIMIT_NOFILE)

if __name__ == "__main__":
    print("nofile:", raise_nofile_limit())
    asyncio.run(main())

Measured with the soft limit raised to the hard limit of 524,288: all 2,000 clients served, p50 177 ms and p99 216 ms, no accept errors, and 2,008 descriptors in use at the peak. If the hard limit itself is too low, raise it where the process is started: LimitNOFILE= in a systemd unit, --ulimit nofile=65536:65536 for docker run, or the container runtime's defaults for Kubernetes pods. Descriptors are cheap in the kernel; memory per connection, measured in load testing WebSocket servers, is the real constraint once the limit is out of the way.

Verify: the start-up log shows the raised soft limit, and a load test above the old limit runs without accept errors.

4. Cap connections below the limit

A high limit is not a plan for running out. Keep a margin of descriptors for the service's own needs by refusing new client connections above a cap:

MAX_CLIENTS = 900                       # limit minus pools, files, and headroom
active = 0

async def handler(reader, writer):
    global active
    if active >= MAX_CLIENTS:
        writer.close()                  # reject now, while we still have descriptors
        return
    active += 1
    try:
        await serve_client(reader, writer)
    finally:
        active -= 1
        writer.close()

Measured with the soft limit left at 1,024 and a cap of 900: 900 clients were served with p99 244 ms, the other 1,100 were closed immediately, and the server logged no accept errors and had no file-open failures — the reserve stayed available for everything else. A rejected client can retry against another instance; a server that hit EMFILE affected every client and every internal operation. Behind a load balancer, report readiness as false while at the cap so new connections go elsewhere, as in implementing health and readiness probes for asyncio.

Verify: at the cap, new clients are rejected promptly and internal operations that open descriptors still succeed.

Descriptor budget for one process A flow of 5 stages. Descriptor budget for one process Read limits getrlimit at start-up Raise soft up to hard, or set it in systemd / docker Reserve pools, files, DNS, the loop Cap clients reject above the cap Monitor fds in use / soft limit The cap keeps EMFILE from ever reaching accept().

5. Monitor descriptors in use

Descriptor exhaustion is easy to predict from one number: descriptors in use as a fraction of the soft limit. Export it, and alert well before 100%:

async def report_fds(interval: float = 10.0):
    soft, _ = resource.getrlimit(resource.RLIMIT_NOFILE)
    while True:
        in_use = len(os.listdir("/proc/self/fd"))          # itself needs one descriptor
        FD_IN_USE.set(in_use)
        FD_LIMIT.set(soft)
        await asyncio.sleep(interval)

Note the irony measured in step 2: at the limit, even listing /proc/self/fd failed, and the statistics task died — report from a margin, and make the reporter tolerate OSError. A count that grows steadily while connection counts are flat points at a descriptor leak — sockets or files that are opened and never closed — which the techniques in Memory & Resource Leaks track down.

Verify: dashboards show descriptors in use against the soft limit, with an alert at about 80%.

What does the descriptor situation call for? A decision on What do you see with 4 outcomes. What does the descriptor situation call for? What do you see? soft limit < peak connections setrlimit to hard at start-up 2,000 served, p99 216 ms hard limit too low LimitNOFILE / --ulimit nofile set where the process starts could still reach the limit cap clients below it 0 accept errors fds grow, connections flat find the leak not a limit problem Never let accept() be the place that discovers the limit.

Verification

A service is safe from descriptor exhaustion when:

  • Its soft limit is logged at start-up and exceeds peak connections plus internal needs.
  • The limit is raised in-process or by its supervisor, never left at a 1,024 default by accident.
  • Client connections are capped below the limit, with prompt rejection above the cap.
  • Descriptors in use are monitored against the limit, by a reporter that tolerates EMFILE.

Diagnostic Hook: alert on any log record containing "out of system resource" or Errno 24. At the limit, asyncio repeats the accept error thousands of times per second, so a log-volume spike from a single message is itself the signal — and a log pipeline that rate-limits or drops it may hide the incident entirely.

Pitfalls & edge cases

  • Trusting your shell's ulimit -n. The service inherits its supervisor's limits.
  • Hitting the limit with a large backlog. Measured: 172,072 accept errors in 10 s.
  • Assuming only new clients suffer. Measured: unrelated file opens failed too.
  • Raising the limit without a cap. Memory or downstream limits become the next wall.

Frequently Asked Questions

How do I fix 'Too many open files' in an asyncio server?

Raise the soft RLIMIT_NOFILE at start-up with resource.setrlimit up to the hard limit, raise the hard limit in systemd or the container if needed, and cap client connections below the limit. With the limit raised, 2,000 clients were served with no errors.

What does 'socket.accept() out of system resource' mean?

accept() failed with EMFILE because the process ran out of file descriptors. asyncio logged it 172,072 times in about 10 s with a backlog of 4,096 in testing.

Does hitting the file descriptor limit only affect new connections?

No: every operation that needs a descriptor fails, including opening files, database reconnects and DNS lookups; a background open() failed 12 times during the test.

How many connections can I allow if my limit is 1,024?

Fewer than 1,024, leaving room for pools, files and the loop; a cap of 900 kept accept errors and file failures at zero.