MemoryError
1 · The lesson
readWhat this error means
Python asked the operating system for memory and was refused. There is no message and no detail — just the name of the exception — because by the time it is raised there is not much room left to describe anything.
You will not usually see it on a laptop, where the OS swaps to disk instead. You see it in containers, which have a hard limit and no swap.
When you see it
Traceback (most recent call last): File "report.py", line 6, in <module> rows = f.readlines() MemoryError
Often you do not see it at all. The container hits its limit first and the kernel kills the process outright:
worker exited with code 137
Exit code 137 is 128 + 9 — killed by SIGKILL, the out-of-memory killer. No traceback is produced, which is what makes it confusing: the logs simply stop.
Why it happens
- Reading a whole file into memory.
f.read()orf.readlines()on a 4 GB export needs 4 GB, plus overhead. fetchall()on a large query, which materialises every row before you touch the first one.- Building a list where a generator would do —
[transform(r) for r in rows]holds every result at once. pd.read_csvwithoutchunksize, where a dataframe can take several times the file size in RAM because of dtypes and index overhead.- Accumulating in a loop and never clearing — appending to a list that lives for the whole run.
How to fix it
Stream the file instead of loading it. A file object is already an iterator over lines, so this holds one line at a time regardless of file size:
total = 0 with open("huge.csv", encoding="utf-8") as f: next(f) # skip the header for line in f: # one line in memory, not the file total += float(line.split(",")[3])
Iterate the cursor rather than fetchall():
with conn.cursor(name="export") as cur: # a named cursor streams server-side cur.execute("SELECT id, email FROM users") for row in cur: write(row)
setup added so this can run · defines conn, write
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) conn = _AutoMock('conn') def write(*_a, **_kw): print('-> write() called') return _AutoMock('write()')
Use a generator when the results are consumed once:
def cleaned(rows): for r in rows: # yields; nothing accumulates yield r.strip().lower()
Chunk the dataframe:
import pandas as pd total = 0 for chunk in pd.read_csv("huge.csv", chunksize=100_000): total += chunk["amount"].sum()
Measure before guessing. tracemalloc is in the standard library and shows which lines are holding memory:
import tracemalloc tracemalloc.start() run_the_job() current, peak = tracemalloc.get_traced_memory() print(f"current {current / 1e6:.1f} MB, peak {peak / 1e6:.1f} MB") for stat in tracemalloc.take_snapshot().statistics("lineno")[:5]: print(stat)
setup added so this can run · defines run_the_job
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def run_the_job(*_a, **_kw): print('-> run_the_job() called') return _AutoMock('run_the_job()')
Raising the container limit is a last resort. If usage grows with input size, a bigger limit only moves the failure to a bigger input.
When you'd actually see this in real code
- A nightly export that worked for a year and died when the table crossed a size threshold.
- A worker killed with 137 and no traceback, which reads like a crash until you check the memory limit.
- An endpoint that loads a whole upload into memory — fine in testing, fatal when several users upload at once.
- A pandas job sized for the sample file, run against the full extract.
Related errors
- [OSError: [Errno 24] Too many open files](../error-too-many-open-files/) — the same shape, a different exhausted resource.
- RecursionError: maximum recursion depth exceeded — stack rather than heap.
See Also
- All Python errors — the full index, by type and by when it happens.
- File I/O, Pathlib, and Large Files — streaming instead of loading.
- Generators — the tool that fixes most of these.
- Performance