PythonMastery
In the wild · CPython

Pre-fork Python servers: how Python stopped wasting their shared memory

Pre-fork servers load an app once and fork workers from it. Python itself kept breaking the memory those workers share; here is what changed in 3.7 and 3.12.

The problem

Large Python web deployments often start one process, load the whole application into it, and then fork worker processes from it. PEP 683 describes why: "For some applications it makes sense to get the application into a desired initial state and then fork the process for each worker. This can result in a large performance improvement, especially memory usage." The same PEP notes: "Several enterprise Python users (e.g. Instagram, YouTube) have taken advantage of this."

The saving comes from copy-on-write: after a fork, parent and child share the same memory pages until one of them writes to a page. The catch is that CPython writes to objects even when your code only reads them. It updates reference counts, and the garbage collector touches its bookkeeping fields. Each write copies a page, and the sharing slowly disappears. PEP 683 says these writes "drastically reduce the benefits and have led to some sub-optimal workarounds."

What changed in Python

The part any team can copy

If you run a pre-fork server, build the expensive shared state once, in the parent, then freeze it. This site can't fork inside a browser tab, but you can see the calls work:

python
import gc

gc.disable()                                  # parent: no collections while loading
cache = {str(i): i * i for i in range(1000)}  # shared state, built once
gc.freeze()                                   # right before fork() in a real server
print(gc.get_freeze_count() > 0)

gc.unfreeze()
print(gc.get_freeze_count())
gc.enable()                                   # children turn collection back on
output
True
0

The general lesson is older than either feature: in a pre-fork server, what you load before the fork is shared almost for free, and what each worker builds after it is paid for once per worker.

the tipTake this away

In a pre-fork server, build shared state once in the parent, then gc.freeze() it before forking. Learn it properly: Performance: Profile, Then Optimise.

Sources

every claim above comes from these