PythonMastery
advanced 24 min read · lesson 5 of 9 in Python Advanced

Multiprocessing & Threading: Beating the GIL

1 · The lesson

read

Async handles concurrency for I/O-bound work — one thread, many awaits. But what about the other two cases? A library that only has a blocking API and won't be rewritten. A numerical kernel that needs every core on your laptop. Async helps with neither.

Python ships two heavier-weight concurrency primitives: threads (multiple flows of execution in one process, sharing memory) and processes (multiple OS processes, each with their own memory). Both have a single API surface — concurrent.futures — that makes switching between them a one-line change. The hard part is knowing which to pick.

The answer starts with the GIL.


1. The GIL — One Line, Half the Battle

CPython's Global Interpreter Lock is a mutex inside the interpreter. At any instant, exactly one thread is executing Python bytecode. Threads still exist, still run, and still help — but only when they're blocked on I/O or running C extension code that releases the GIL (NumPy, hashlib, file I/O, network calls).

The consequences:

  • Threads for I/O — yes. A thread waiting on a socket has released the GIL; another thread runs.
  • Threads for CPU — no. A pure-Python loop hogs the GIL. Ten threads × ten cores still gives you one core's worth of throughput. Sometimes less, after the lock-handover overhead.
  • Processes for CPU — yes. Each Python process has its own GIL. Eight processes can use eight cores.

3.13 note: PEP 703 added an experimental free-threaded build of CPython (no GIL). It's opt-in, requires a separate interpreter binary, and most C extensions don't support it yet. For everything you'll write in the next year or two, the GIL is real and the decisions below stand.


2. The Decision Matrix (Memorise This)

WorkloadUse
Many I/O calls, async-native libraries existasyncio
I/O-bound, only blocking libraries availableThreadPoolExecutor
CPU-bound (pure Python: math, parsing, regex)ProcessPoolExecutor
CPU-bound (heavy NumPy/Numba/Cython)Threads work — those release GIL
Shared memory required, simple coordinationThreads + Lock
Untrusted code, isolation desiredProcesses

Three trip-wires for the wrong choice:

1. "I added threads and the CPU work isn't any faster" → GIL. Use processes.
2. "I added processes and now my objects don't share state" → processes have separate memory. Use a Manager, Queue, or shared_memory.
3. "I added async and the requests library still blocks everything" → requests is sync. Use httpx.AsyncClient, or run requests through asyncio.to_thread.


3. threading.Thread — The Bare Metal

The low-level API. You'll mostly use ThreadPoolExecutor (Section 7), but the primitives are worth seeing once.

python
import threading
import time

def worker(name, delay):
    print(f"{name}: starting")
    time.sleep(delay)
    print(f"{name}: done")

t1 = threading.Thread(target=worker, args=("a", 1.0))
t2 = threading.Thread(target=worker, args=("b", 0.5))

t1.start(); t2.start()
t1.join(); t2.join()                         # wait for both to finish
print("all done")

start() schedules the thread; the OS decides when it actually runs. join() blocks until the thread finishes. Without join, the main thread can exit while workers are still mid-flight — usually fine, occasionally a disaster.

Daemon vs non-daemon: a daemon thread (Thread(target=..., daemon=True)) is killed when the main thread exits. Non-daemon threads keep the program alive until they finish. Background "tick every minute" workers: daemon. Threads doing real work the program depends on: non-daemon.


4. Locks, RLocks, Semaphores

Multiple threads writing to the same data structure without synchronisation is a recipe for corrupted state. Locks fix it.

python
import threading

counter = 0
lock = threading.Lock()

def increment(n):
    global counter
    for _ in range(n):
        with lock:                           # only one thread inside at a time
            counter += 1                     # the unsafe critical section

threads = [threading.Thread(target=increment, args=(100_000,)) for _ in range(8)]
for t in threads: t.start()
for t in threads: t.join()

print(counter)                               # 800_000 — without the lock, you'd see a smaller, random number

The primitives:

  • Lock — mutual exclusion. One holder at a time. Cannot be re-acquired by the same thread (deadlock).
  • RLock — re-entrant lock. The same thread can acquire multiple times (releases must match). Use when a function holding the lock calls another that also needs it.
  • Semaphore(n) — allow up to n simultaneous holders. The threaded equivalent of asyncio.Semaphore. Use to cap concurrent HTTP requests, DB connections, file handles.
  • Event — a one-shot flag for "something happened"; threads wait() until another thread set()s it.
  • Condition — a lock plus a wait/notify protocol. Producer/consumer signalling. Almost always replaced by queue.Queue (next section) — easier and correct by default.

with lock: is non-negotiable. Forgetting to release a lock — through an early return, an exception, anything — deadlocks every other thread.


5. queue.Queue — Thread-Safe Data Passing

Locks are easy to use wrong. A queue is hard to use wrong. For "one thread produces, another consumes," reach for queue.Queue first.

python
import queue, threading, time

q = queue.Queue(maxsize=10)

def producer():
    for i in range(20):
        q.put(i)                             # blocks if queue is full
        print(f"produced {i}")
    q.put(None)                              # sentinel: tells consumer to stop

def consumer():
    while True:
        item = q.get()                       # blocks until something's there
        if item is None:
            q.task_done()
            return
        print(f"  consumed {item}")
        q.task_done()

threading.Thread(target=producer).start()
threading.Thread(target=consumer).start()

Queue is thread-safe: every put and get is atomic. The producer can hammer it from one thread while the consumer drains it from another and nothing corrupts. The maxsize cap gives you backpressure — fast producers can't outrun slow consumers and OOM the program.

Variants: LifoQueue (stack), PriorityQueue (priority-ordered tuples). Same API.


6. Thread-Local Data

Sometimes you want each thread to have its own copy of a "global" — a database connection, a request context, a per-thread cache. threading.local() is the storage.

python
import threading

ctx = threading.local()

def init_thread(name):
    ctx.user = name                          # only this thread sees it

def use_thread():
    print(f"hello, {ctx.user}")              # works only after init_thread on same thread

threading.Thread(target=lambda: (init_thread("alice"), use_thread())).start()
threading.Thread(target=lambda: (init_thread("bob"), use_thread())).start()

Each thread gets its own attribute namespace on the local() object. Read it on a thread that hasn't set it and you get AttributeError. Frameworks like Flask use this for the per-request context — but more modern code prefers contextvars (which also works with async).


7. ThreadPoolExecutor — The Preferred API

You almost never want to manually create Thread objects in modern code. concurrent.futures.ThreadPoolExecutor is the high-level pool: submit work, get futures back, let the pool manage the threads.

python
from concurrent.futures import ThreadPoolExecutor, as_completed
import time

def fetch(url):
    time.sleep(0.5)                          # pretend this is a network call
    return f"got {url}"

urls = [f"https://x/{i}" for i in range(20)]

with ThreadPoolExecutor(max_workers=8) as pool:
    # Option A: map — results in input order
    for result in pool.map(fetch, urls):
        print(result)

    # Option B: submit + as_completed — process in completion order
    futures = {pool.submit(fetch, u): u for u in urls}
    for fut in as_completed(futures):
        url = futures[fut]
        try:
            print(url, "->", fut.result())
        except Exception as e:
            print(url, "FAILED:", e)

Three workhorses:

  • pool.submit(fn, *args) — returns a Future. Call .result() to block until done (or raise the function's exception). The building block.
  • pool.map(fn, iterable) — lazy iterator over results in input order. Convenient when order matters and you don't need per-item error handling.
  • as_completed(futures) — yields futures as they finish. Best when items take different times and you want to start handling fast ones immediately.

The with block joins all workers before exiting — no orphan threads.


8. ProcessPoolExecutor — Same Shape, Different Universe

python
from concurrent.futures import ProcessPoolExecutor

def heavy(n):
    # pure-Python CPU loop — would not parallelise with threads
    return sum(i * i for i in range(n))

if __name__ == "__main__":
    with ProcessPoolExecutor(max_workers=4) as pool:
        results = list(pool.map(heavy, [10_000_000] * 8))
    print(results)

The API is deliberately identical to ThreadPoolExecutor. Swap Thread → Process and you switch from "share memory, GIL-bound" to "separate processes, true parallelism." The cost is the serialisation tax: arguments and return values are pickled to cross the process boundary.

if __name__ == "__main__": is required on Windows. We'll explain why in Section 11.


9. multiprocessing.Pool — The Older API

Before concurrent.futures (Python 3.2), multiprocessing.Pool was the way. It's still in active use, and you'll see it in old codebases. The shape is similar:

python
from multiprocessing import Pool

if __name__ == "__main__":
    with Pool(processes=4) as p:
        print(p.map(heavy, [10_000_000] * 8))
        print(p.apply(heavy, (1_000_000,)))                  # one-shot, blocking
        print(p.apply_async(heavy, (1_000_000,)).get())      # one-shot, async
        # imap_unordered: stream results as they finish, like as_completed
        for r in p.imap_unordered(heavy, [10_000_000] * 8):
            print(r)
+ setup added so this can run · defines heavy
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

heavy = _AutoMock('heavy')

Reasons you might still use it:

  • imap_unordered with a chunksize is often the fastest way to stream a big batch through a process pool.
  • Existing code uses it and rewriting isn't justified.

For new code, prefer ProcessPoolExecutor — same capability, cleaner ergonomics, composes with asyncio.run_in_executor.


10. Sharing Data Between Processes

Threads share memory for free. Processes do not. Three patterns to bridge them:

multiprocessing.Queue — message passing

python
from multiprocessing import Process, Queue

def worker(q):
    q.put("hello from worker")

if __name__ == "__main__":
    q = Queue()
    p = Process(target=worker, args=(q,))
    p.start()
    print(q.get())                           # "hello from worker"
    p.join()

Different from queue.Queue — this one is process-safe (uses a pipe + locks under the hood). Items are pickled on put and unpickled on get.

multiprocessing.Manager — proxied mutable state

python
from multiprocessing import Manager, Process

def append(shared_list, item):
    shared_list.append(item)

if __name__ == "__main__":
    with Manager() as m:
        shared = m.list()                    # proxy to a list living in a manager process
        ps = [Process(target=append, args=(shared, i)) for i in range(5)]
        for p in ps: p.start()
        for p in ps: p.join()
        print(list(shared))                  # [0, 1, 2, 3, 4] (some order)

A Manager spawns a separate server process that owns the actual objects; everyone else gets proxies. Every method call serialises arguments, ships them over a socket, and ships results back. Convenient. Slow. Don't put hot-path objects in there.

multiprocessing.shared_memory — zero-copy bytes (3.8+)

python
from multiprocessing import shared_memory, Process
import numpy as np

def increment(name, shape, dtype):
    shm = shared_memory.SharedMemory(name=name)
    arr = np.ndarray(shape, dtype=dtype, buffer=shm.buf)
    arr += 1
    shm.close()

if __name__ == "__main__":
    arr = np.zeros((1_000_000,), dtype=np.int64)
    shm = shared_memory.SharedMemory(create=True, size=arr.nbytes)
    backed = np.ndarray(arr.shape, dtype=arr.dtype, buffer=shm.buf)
    backed[:] = arr

    procs = [Process(target=increment, args=(shm.name, arr.shape, arr.dtype))
             for _ in range(4)]
    for p in procs: p.start()
    for p in procs: p.join()

    print(backed[:5])                        # [4, 4, 4, 4, 4]
    shm.close()
    shm.unlink()                             # free the shared memory block

A raw memory block visible to every process by name. Combined with NumPy, this is how you pass gigabytes of data between processes for the cost of an mmap — no pickling, no copying. The fastest option, and the one most likely to corrupt itself if you don't add your own locks. Reach for it when you've already measured pickling as the bottleneck.


11. if __name__ == "__main__": and the Windows Spawn Trap

python
from multiprocessing import Process

def hi(): print("hi from child")

# DANGER on Windows
p = Process(target=hi)
p.start()

On macOS (since 3.8) and Windows, multiprocessing uses the spawn start method: the child process launches a fresh Python interpreter and re-imports your script. If the top level of your script also spawns processes — boom, infinite recursion of subprocess launches. Your CPU pegs and your task manager fills with python.exe.

The fix is the boilerplate every multiprocessing example has:

python
if __name__ == "__main__":
    p = Process(target=hi)
    p.start()
    p.join()
+ setup added so this can run · defines Process, hi
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def Process(*_a, **_kw):
    print('-> Process() called')
    return _AutoMock('Process()')
hi = _AutoMock('hi')

When the child re-imports the script, __name__ is "__mp_main__", not "__main__", so the block doesn't run again. Everything inside that guard is "code that runs only in the parent."

Linux uses fork by default — no re-import, no problem — but writing portable code means following the spawn rules everywhere. The guard costs nothing.


12. Map-Reduce Pattern

The bread-and-butter use of a process pool: split work, parallelise the map, aggregate the result.

python
from concurrent.futures import ProcessPoolExecutor

def word_counts(text_chunk):
    counts = {}
    for word in text_chunk.split():
        counts[word] = counts.get(word, 0) + 1
    return counts

def merge(a, b):
    out = dict(a)
    for k, v in b.items():
        out[k] = out.get(k, 0) + v
    return out

def parallel_word_count(chunks):
    with ProcessPoolExecutor() as pool:
        per_chunk = pool.map(word_counts, chunks)
    total = {}
    for d in per_chunk:
        total = merge(total, d)
    return total

if __name__ == "__main__":
    chunks = ["the quick brown fox"] * 8
    print(parallel_word_count(chunks))

Map runs in parallel across workers. Reduce (merge) runs on the main process — typically fast, since each map result is already aggregated. The shape is identical to MapReduce/Spark; the only thing missing is the cluster.


13. A Benchmark Worth Internalising

The canonical demo: a CPU-bound function, run serially vs through a thread pool vs through a process pool.

python
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor
import time

def heavy(n):
    return sum(i * i for i in range(n))

WORK = [10_000_000] * 8

def serial():
    return [heavy(n) for n in WORK]

def threads():
    with ThreadPoolExecutor(max_workers=8) as pool:
        return list(pool.map(heavy, WORK))

def processes():
    with ProcessPoolExecutor(max_workers=8) as pool:
        return list(pool.map(heavy, WORK))

if __name__ == "__main__":
    for label, fn in [("serial", serial), ("threads", threads), ("processes", processes)]:
        t0 = time.perf_counter()
        fn()
        print(f"{label:>10}: {time.perf_counter()-t0:.2f}s")

On a typical 8-core laptop:

python
    serial: ~6.0s
   threads: ~6.0s   (no faster — GIL serialises the bytecode loop)
 processes: ~0.9s   (~6.7× speedup — bounded by core count and pickling overhead)

Threads gave you zero. Processes gave you near-linear scaling. That gap is the GIL. If your workload looks like this, multiprocessing is the only standard-library answer.

(If heavy were numpy.sum(numpy.arange(n)**2) instead — the loop lives in C and releases the GIL — threads would scale too. Knowing which path your code takes is half the optimisation work.)


Common Mistakes

1. Threading CPU-bound code and expecting speedup

The most common misconception. Pure-Python loops do not parallelise with threads, full stop. Profile first; if the hot loop is Python bytecode, switch to processes or move the loop into NumPy/C.

2. Forgetting if __name__ == "__main__": on Windows/macOS

Endless subprocess spawning, CPU pegged, task manager full of python.exe. The guard is mandatory for any module that calls Process, Pool, or ProcessPoolExecutor.

3. Mutating "shared" state in subprocesses

python
counter = 0
def bump(): global counter; counter += 1

with ProcessPoolExecutor() as p:
    list(p.map(bump, range(10)))

print(counter)                               # 0 — each child has its own `counter`
+ setup added so this can run · defines ProcessPoolExecutor
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def ProcessPoolExecutor(*_a, **_kw):
    print('-> ProcessPoolExecutor() called')
    return _AutoMock('ProcessPoolExecutor()')

Each subprocess gets a copy of the parent's state at spawn time. Mutations don't propagate back. Use Manager, Queue, or shared_memory — or design out shared state by returning values from workers and aggregating in the parent.

4. Deadlocks from inconsistent lock ordering

Thread 1 holds A, wants B. Thread 2 holds B, wants A. Both wait forever. The fix is discipline: pick a global order (e.g. by id(lock)) and always acquire in that order. Better: design out the need for two locks at once.

5. Pickle errors on closures, lambdas, local functions

python
def outer():
    def inner(x): return x * 2
    with ProcessPoolExecutor() as p:
        list(p.map(inner, range(10)))        # PicklingError: can't pickle local function
+ setup added so this can run · defines ProcessPoolExecutor
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def ProcessPoolExecutor(*_a, **_kw):
    print('-> ProcessPoolExecutor() called')
    return _AutoMock('ProcessPoolExecutor()')

Process pools pickle the function to ship it to workers. Pickling needs the function to be importable at module top level. Move inner out, or use multiprocess (third-party fork using dill) — but module-level functions are the right answer.

6. Pool overhead dwarfing the work

Submitting 100,000 tiny tasks to a process pool spends more time pickling than computing. Two fixes: batch the work (each task does 1000 items instead of 1) or pass chunksize=N to pool.map so the pool batches for you.

7. Not joining/closing the pool

Always with ThreadPoolExecutor(...) as pool:. The with block calls shutdown(wait=True) on exit — workers drain, threads stop. Without it, workers can outlive the main thread and prevent the program from exiting.


🎯 Your Turn — parallel_map(fn, items, mode="thread")

Write a function that parallelises a map across either threads or processes, picked by a mode argument. Then benchmark it on a CPU-bound function and prove to yourself that only process mode actually speeds things up.

Requirements:

  • Signature: parallel_map(fn, items, mode="thread", max_workers=None).
  • mode="thread" uses ThreadPoolExecutor; mode="process" uses ProcessPoolExecutor; anything else raises ValueError.
  • Returns results in input order (same as pool.map).
  • If max_workers is None, let the executor choose (it defaults to a reasonable value).
  • Include a benchmark block (under if __name__ == "__main__":) that times a CPU-bound function in serial, thread, and process mode.
python
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor

def parallel_map(fn, items, mode="thread", max_workers=None):
    # TODO 1: pick the executor class based on `mode`
    # TODO 2: open the executor in a `with` block
    # TODO 3: return list(executor.map(fn, items))
    ...

def cpu_heavy(n):
    return sum(i * i for i in range(n))

if __name__ == "__main__":
    work = [2_000_000] * 8
    # TODO 4: time serial baseline
    # TODO 5: time parallel_map(cpu_heavy, work, mode="thread")
    # TODO 6: time parallel_map(cpu_heavy, work, mode="process")
    # Print the three timings.
Hint 1 — Mapping mode to executor A simple dict works: {"thread": ThreadPoolExecutor, "process": ProcessPoolExecutor}[mode]. Use .get(mode) and raise ValueError if the result is None. Or use a plain if/elif/else chain.
Hint 2 — Why the benchmark needs the __main__ guard ProcessPoolExecutor spawns subprocesses that re-import your script. Without the guard, those subprocesses would also try to run the benchmark, which would spawn more subprocesses, which would... your machine catches fire. Every multiprocessing example you write must live under if __name__ == "__main__":.
Show full solution
python
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor
import time


def parallel_map(fn, items, mode="thread", max_workers=None):
    """Parallelise `fn` over `items` using threads or processes.

    Results returned in input order, like the built-in `map`.
    """
    executor_cls = {
        "thread": ThreadPoolExecutor,
        "process": ProcessPoolExecutor,
    }.get(mode)
    if executor_cls is None:
        raise ValueError(f"mode must be 'thread' or 'process', got {mode!r}")

    with executor_cls(max_workers=max_workers) as pool:
        return list(pool.map(fn, items))


def cpu_heavy(n):
    """Pure-Python CPU loop — does NOT release the GIL."""
    return sum(i * i for i in range(n))


def timed(label, fn):
    t0 = time.perf_counter()
    fn()
    elapsed = time.perf_counter() - t0
    print(f"{label:>14}: {elapsed:.2f}s")


if __name__ == "__main__":
    work = [2_000_000] * 8

    timed("serial",       lambda: [cpu_heavy(n) for n in work])
    timed("thread pool",  lambda: parallel_map(cpu_heavy, work, mode="thread"))
    timed("process pool", lambda: parallel_map(cpu_heavy, work, mode="process"))

Typical output on an 8-core machine:

python
        serial: 1.20s
   thread pool: 1.22s    <- no speedup; the GIL serialises bytecode
  process pool: 0.18s    <- ~6.5x speedup; each process owns its own GIL

What this benchmark proves, in one screenful:

  • Threads are useless here. Eight workers, one core's worth of throughput. The GIL forces the eight Python interpreters to take turns on the bytecode loop.
  • Processes scale. Each worker is a separate interpreter with its own GIL, running on its own core. Speedup is sub-linear (pickling argument and result has overhead, and the OS doesn't give you all 8 cores at 100%) but the order of magnitude is right.
  • The API is identical. Same parallel_map function, same pool.map, same with block. The one-word mode change is all it takes to switch concurrency models.

A more interesting variant: replace cpu_heavy with numpy.sum(numpy.arange(n) ** 2) — that loop runs in C and releases the GIL. Threads will then scale just as well as processes, and with no pickling overhead. The general rule: GIL-releasing work parallelises with threads; pure-Python work needs processes.

For real workloads, profile first. The right knob to turn is often "vectorise with NumPy" or "rewrite the hot loop in Cython," not "throw more workers at it." Concurrency is the answer when the work is already efficient and you simply need more of it at once.


What You Learned

  • The GIL — one thread of Python bytecode at a time per process. Threads help with I/O and GIL-releasing C code; they don't help with pure-Python CPU work.
  • Decision matrix: I/O + async libs → asyncio; I/O + blocking libs → threads; CPU-bound → processes.
  • threading.Thread, Lock, RLock, Semaphore — the low-level primitives. Always acquire under with.
  • queue.Queue — thread-safe producer/consumer. Almost always better than rolling your own with locks.
  • threading.local() — per-thread state. contextvars is the async-friendly modern alternative.
  • concurrent.futures.ThreadPoolExecutor — the modern preferred thread API. submit returns a Future; map preserves order; as_completed yields in completion order.
  • concurrent.futures.ProcessPoolExecutor — same shape, separate processes, true parallelism. Arguments and results are pickled.
  • multiprocessing.Pool — older API; imap_unordered is still useful. New code: prefer ProcessPoolExecutor.
  • Sharing across processes: Queue (messages), Manager (proxied state — slow), shared_memory (zero-copy bytes — fast, dangerous).
  • if __name__ == "__main__": is mandatory with multiprocessing on Windows/macOS (spawn) — otherwise child processes re-import the script and respawn forever.
  • Map-reduce — pool.map(map_fn, chunks) + a serial reduce. The shape of every batch job you'll ever write.
  • Pickling tax — function arguments and return values cross process boundaries serialised. Batch small tasks (chunksize=) and avoid large per-call payloads.
  • Pool overhead can exceed benefit for tiny tasks. Measure before parallelising.

Next: Type Hints in Depth — Protocol, TypeVar, Generic, Self, ParamSpec, and how the modern Python type system actually catches bugs at the boundary.