OSError: [Errno 24] Too many open files
1 · The lesson
readWhat this error means
Your process hit the operating system's limit on open file descriptors. A descriptor is not only a file — every socket, pipe and database connection is one too — so a program that never opens a file can still exhaust the limit through network connections alone.
The error almost always means a leak. Something is being opened in a loop and never closed.
When you see it
Traceback (most recent call last): File "worker.py", line 31, in fetch_all r = requests.get(url) OSError: [Errno 24] Too many open files
It rarely appears where the bug is. The leak accumulates quietly, and the error surfaces at whichever innocent line happened to ask for the descriptor that crossed the limit — often minutes or hours into a run.
Why it happens
- Files opened without
with.f = open(path)leaves the descriptor alive until the object is garbage collected, which may be much later, or never in a reference cycle. - A new
requests.Sessionper call. Each session holds its own connection pool. Creating one per request in a loop leaks sockets steadily. - Sockets left in
CLOSE_WAITbecause the code stopped reading but never closed. - A database connection per request instead of a pool, with connections never returned.
- A limit that is simply too low for the workload. The default is often 1024, which a service holding a few hundred concurrent connections can reach legitimately.
How to fix it
Count them, and watch the number move. This is the fastest way to tell a leak from a limit:
ls /proc/$(pgrep -f worker.py)/fd | wc -l # Linux lsof -p $(pgrep -f worker.py) | wc -l # macOS
Run it a few times a minute apart. Climbing steadily means a leak; flat but high means the limit is genuinely too small.
From inside the process:
import resource soft, hard = resource.getrlimit(resource.RLIMIT_NOFILE) print(f"file descriptor limit: soft={soft} hard={hard}")
Always use with for files — it closes on the way out, including when an exception is raised:
with open("data.csv", encoding="utf-8") as f: rows = f.readlines() # closed here, guaranteed
Reuse one session for many requests. This is the single most common source of the leak in Python services:
import requests session = requests.Session() # once, at module or app level for url in urls: r = session.get(url, timeout=(3.05, 10))
setup added so this can run · defines urls
urls = ["alpha", "beta", "gamma"]
Pool your database connections rather than opening one per request, and return them:
from psycopg_pool import ConnectionPool pool = ConnectionPool("postgresql://db/shop", min_size=2, max_size=10) with pool.connection() as conn: # borrowed here, returned at the end of the block conn.execute("SELECT 1")
Raise the limit only once you know it is not a leak. Raising it on a leaking process buys minutes and hides the cause:
# docker-compose.yml
services:
worker:
ulimits:
nofile: { soft: 65535, hard: 65535 }When you'd actually see this in real code
- A crawler creating a
Sessioninside the loop instead of outside it — works for ten URLs, dies at ten thousand. - A log handler opened per message rather than once at startup.
- An image pipeline using
Image.open()withoutwith, holding every file until garbage collection catches up. - A service that runs fine for days and falls over on the busiest afternoon, because the leak needed enough traffic to reach the ceiling.
Related errors
- MemoryError — the same shape of bug, a different exhausted resource.
- [ConnectionRefusedError: [Errno 111] Connection refused](../error-connection-refused/) — what clients see once you can no longer accept connections.
- [PermissionError: [Errno 13] Permission denied](../error-permissionerror/)
See Also
- All Python errors — the full index, by type and by when it happens.
- File I/O, Pathlib, and Large Files — why
withexists. - Databases in production — pooling.