Multiprocessing & Threading: Beating the GIL
1 · The lesson
readAsync handles concurrency for I/O-bound work — one thread, many awaits. But what about the other two cases? A library that only has a blocking API and won't be rewritten. A numerical kernel that needs every core on your laptop. Async helps with neither.
Python ships two heavier-weight concurrency primitives: threads (multiple flows of execution in one process, sharing memory) and processes (multiple OS processes, each with their own memory). Both have a single API surface — concurrent.futures — that makes switching between them a one-line change. The hard part is knowing which to pick.
The answer starts with the GIL.
1. The GIL — One Line, Half the Battle
CPython's Global Interpreter Lock is a mutex inside the interpreter. At any instant, exactly one thread is executing Python bytecode. Threads still exist, still run, and still help — but only when they're blocked on I/O or running C extension code that releases the GIL (NumPy, hashlib, file I/O, network calls).
The consequences:
- Threads for I/O — yes. A thread waiting on a socket has released the GIL; another thread runs.
- Threads for CPU — no. A pure-Python loop hogs the GIL. Ten threads × ten cores still gives you one core's worth of throughput. Sometimes less, after the lock-handover overhead.
- Processes for CPU — yes. Each Python process has its own GIL. Eight processes can use eight cores.
3.13 note: PEP 703 added an experimental free-threaded build of CPython (no GIL). It's opt-in, requires a separate interpreter binary, and most C extensions don't support it yet. For everything you'll write in the next year or two, the GIL is real and the decisions below stand.
2. The Decision Matrix (Memorise This)
| Workload | Use |
|---|---|
| Many I/O calls, async-native libraries exist | asyncio |
| I/O-bound, only blocking libraries available | ThreadPoolExecutor |
| CPU-bound (pure Python: math, parsing, regex) | ProcessPoolExecutor |
| CPU-bound (heavy NumPy/Numba/Cython) | Threads work — those release GIL |
| Shared memory required, simple coordination | Threads + Lock |
| Untrusted code, isolation desired | Processes |
Three trip-wires for the wrong choice:
1. "I added threads and the CPU work isn't any faster" → GIL. Use processes.
2. "I added processes and now my objects don't share state" → processes have separate memory. Use a Manager, Queue, or shared_memory.
3. "I added async and the requests library still blocks everything" → requests is sync. Use httpx.AsyncClient, or run requests through asyncio.to_thread.
3. threading.Thread — The Bare Metal
The low-level API. You'll mostly use ThreadPoolExecutor (Section 7), but the primitives are worth seeing once.
import threading import time def worker(name, delay): print(f"{name}: starting") time.sleep(delay) print(f"{name}: done") t1 = threading.Thread(target=worker, args=("a", 1.0)) t2 = threading.Thread(target=worker, args=("b", 0.5)) t1.start(); t2.start() t1.join(); t2.join() # wait for both to finish print("all done")
start() schedules the thread; the OS decides when it actually runs. join() blocks until the thread finishes. Without join, the main thread can exit while workers are still mid-flight — usually fine, occasionally a disaster.
Daemon vs non-daemon: a daemon thread (Thread(target=..., daemon=True)) is killed when the main thread exits. Non-daemon threads keep the program alive until they finish. Background "tick every minute" workers: daemon. Threads doing real work the program depends on: non-daemon.
4. Locks, RLocks, Semaphores
Multiple threads writing to the same data structure without synchronisation is a recipe for corrupted state. Locks fix it.
import threading counter = 0 lock = threading.Lock() def increment(n): global counter for _ in range(n): with lock: # only one thread inside at a time counter += 1 # the unsafe critical section threads = [threading.Thread(target=increment, args=(100_000,)) for _ in range(8)] for t in threads: t.start() for t in threads: t.join() print(counter) # 800_000 — without the lock, you'd see a smaller, random number
The primitives:
Lock— mutual exclusion. One holder at a time. Cannot be re-acquired by the same thread (deadlock).RLock— re-entrant lock. The same thread canacquiremultiple times (releases must match). Use when a function holding the lock calls another that also needs it.Semaphore(n)— allow up tonsimultaneous holders. The threaded equivalent ofasyncio.Semaphore. Use to cap concurrent HTTP requests, DB connections, file handles.Event— a one-shot flag for "something happened"; threadswait()until another threadset()s it.Condition— a lock plus a wait/notify protocol. Producer/consumer signalling. Almost always replaced byqueue.Queue(next section) — easier and correct by default.
with lock: is non-negotiable. Forgetting to release a lock — through an early return, an exception, anything — deadlocks every other thread.
5. queue.Queue — Thread-Safe Data Passing
Locks are easy to use wrong. A queue is hard to use wrong. For "one thread produces, another consumes," reach for queue.Queue first.
import queue, threading, time q = queue.Queue(maxsize=10) def producer(): for i in range(20): q.put(i) # blocks if queue is full print(f"produced {i}") q.put(None) # sentinel: tells consumer to stop def consumer(): while True: item = q.get() # blocks until something's there if item is None: q.task_done() return print(f" consumed {item}") q.task_done() threading.Thread(target=producer).start() threading.Thread(target=consumer).start()
Queue is thread-safe: every put and get is atomic. The producer can hammer it from one thread while the consumer drains it from another and nothing corrupts. The maxsize cap gives you backpressure — fast producers can't outrun slow consumers and OOM the program.
Variants: LifoQueue (stack), PriorityQueue (priority-ordered tuples). Same API.
6. Thread-Local Data
Sometimes you want each thread to have its own copy of a "global" — a database connection, a request context, a per-thread cache. threading.local() is the storage.
import threading ctx = threading.local() def init_thread(name): ctx.user = name # only this thread sees it def use_thread(): print(f"hello, {ctx.user}") # works only after init_thread on same thread threading.Thread(target=lambda: (init_thread("alice"), use_thread())).start() threading.Thread(target=lambda: (init_thread("bob"), use_thread())).start()
Each thread gets its own attribute namespace on the local() object. Read it on a thread that hasn't set it and you get AttributeError. Frameworks like Flask use this for the per-request context — but more modern code prefers contextvars (which also works with async).
7. ThreadPoolExecutor — The Preferred API
You almost never want to manually create Thread objects in modern code. concurrent.futures.ThreadPoolExecutor is the high-level pool: submit work, get futures back, let the pool manage the threads.
from concurrent.futures import ThreadPoolExecutor, as_completed import time def fetch(url): time.sleep(0.5) # pretend this is a network call return f"got {url}" urls = [f"https://x/{i}" for i in range(20)] with ThreadPoolExecutor(max_workers=8) as pool: # Option A: map — results in input order for result in pool.map(fetch, urls): print(result) # Option B: submit + as_completed — process in completion order futures = {pool.submit(fetch, u): u for u in urls} for fut in as_completed(futures): url = futures[fut] try: print(url, "->", fut.result()) except Exception as e: print(url, "FAILED:", e)
Three workhorses:
pool.submit(fn, *args)— returns aFuture. Call.result()to block until done (or raise the function's exception). The building block.pool.map(fn, iterable)— lazy iterator over results in input order. Convenient when order matters and you don't need per-item error handling.as_completed(futures)— yields futures as they finish. Best when items take different times and you want to start handling fast ones immediately.
The with block joins all workers before exiting — no orphan threads.
8. ProcessPoolExecutor — Same Shape, Different Universe
from concurrent.futures import ProcessPoolExecutor def heavy(n): # pure-Python CPU loop — would not parallelise with threads return sum(i * i for i in range(n)) if __name__ == "__main__": with ProcessPoolExecutor(max_workers=4) as pool: results = list(pool.map(heavy, [10_000_000] * 8)) print(results)
The API is deliberately identical to ThreadPoolExecutor. Swap Thread → Process and you switch from "share memory, GIL-bound" to "separate processes, true parallelism." The cost is the serialisation tax: arguments and return values are pickled to cross the process boundary.
if __name__ == "__main__": is required on Windows. We'll explain why in Section 11.
9. multiprocessing.Pool — The Older API
Before concurrent.futures (Python 3.2), multiprocessing.Pool was the way. It's still in active use, and you'll see it in old codebases. The shape is similar:
from multiprocessing import Pool if __name__ == "__main__": with Pool(processes=4) as p: print(p.map(heavy, [10_000_000] * 8)) print(p.apply(heavy, (1_000_000,))) # one-shot, blocking print(p.apply_async(heavy, (1_000_000,)).get()) # one-shot, async # imap_unordered: stream results as they finish, like as_completed for r in p.imap_unordered(heavy, [10_000_000] * 8): print(r)
setup added so this can run · defines heavy
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) heavy = _AutoMock('heavy')
Reasons you might still use it:
imap_unorderedwith achunksizeis often the fastest way to stream a big batch through a process pool.- Existing code uses it and rewriting isn't justified.
For new code, prefer ProcessPoolExecutor — same capability, cleaner ergonomics, composes with asyncio.run_in_executor.
10. Sharing Data Between Processes
Threads share memory for free. Processes do not. Three patterns to bridge them:
multiprocessing.Queue — message passing
from multiprocessing import Process, Queue def worker(q): q.put("hello from worker") if __name__ == "__main__": q = Queue() p = Process(target=worker, args=(q,)) p.start() print(q.get()) # "hello from worker" p.join()
Different from queue.Queue — this one is process-safe (uses a pipe + locks under the hood). Items are pickled on put and unpickled on get.
multiprocessing.Manager — proxied mutable state
from multiprocessing import Manager, Process def append(shared_list, item): shared_list.append(item) if __name__ == "__main__": with Manager() as m: shared = m.list() # proxy to a list living in a manager process ps = [Process(target=append, args=(shared, i)) for i in range(5)] for p in ps: p.start() for p in ps: p.join() print(list(shared)) # [0, 1, 2, 3, 4] (some order)
A Manager spawns a separate server process that owns the actual objects; everyone else gets proxies. Every method call serialises arguments, ships them over a socket, and ships results back. Convenient. Slow. Don't put hot-path objects in there.
multiprocessing.shared_memory — zero-copy bytes (3.8+)
from multiprocessing import shared_memory, Process import numpy as np def increment(name, shape, dtype): shm = shared_memory.SharedMemory(name=name) arr = np.ndarray(shape, dtype=dtype, buffer=shm.buf) arr += 1 shm.close() if __name__ == "__main__": arr = np.zeros((1_000_000,), dtype=np.int64) shm = shared_memory.SharedMemory(create=True, size=arr.nbytes) backed = np.ndarray(arr.shape, dtype=arr.dtype, buffer=shm.buf) backed[:] = arr procs = [Process(target=increment, args=(shm.name, arr.shape, arr.dtype)) for _ in range(4)] for p in procs: p.start() for p in procs: p.join() print(backed[:5]) # [4, 4, 4, 4, 4] shm.close() shm.unlink() # free the shared memory block
A raw memory block visible to every process by name. Combined with NumPy, this is how you pass gigabytes of data between processes for the cost of an mmap — no pickling, no copying. The fastest option, and the one most likely to corrupt itself if you don't add your own locks. Reach for it when you've already measured pickling as the bottleneck.
11. if __name__ == "__main__": and the Windows Spawn Trap
from multiprocessing import Process def hi(): print("hi from child") # DANGER on Windows p = Process(target=hi) p.start()
On macOS (since 3.8) and Windows, multiprocessing uses the spawn start method: the child process launches a fresh Python interpreter and re-imports your script. If the top level of your script also spawns processes — boom, infinite recursion of subprocess launches. Your CPU pegs and your task manager fills with python.exe.
The fix is the boilerplate every multiprocessing example has:
if __name__ == "__main__": p = Process(target=hi) p.start() p.join()
setup added so this can run · defines Process, hi
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def Process(*_a, **_kw): print('-> Process() called') return _AutoMock('Process()') hi = _AutoMock('hi')
When the child re-imports the script, __name__ is "__mp_main__", not "__main__", so the block doesn't run again. Everything inside that guard is "code that runs only in the parent."
Linux uses fork by default — no re-import, no problem — but writing portable code means following the spawn rules everywhere. The guard costs nothing.
12. Map-Reduce Pattern
The bread-and-butter use of a process pool: split work, parallelise the map, aggregate the result.
from concurrent.futures import ProcessPoolExecutor def word_counts(text_chunk): counts = {} for word in text_chunk.split(): counts[word] = counts.get(word, 0) + 1 return counts def merge(a, b): out = dict(a) for k, v in b.items(): out[k] = out.get(k, 0) + v return out def parallel_word_count(chunks): with ProcessPoolExecutor() as pool: per_chunk = pool.map(word_counts, chunks) total = {} for d in per_chunk: total = merge(total, d) return total if __name__ == "__main__": chunks = ["the quick brown fox"] * 8 print(parallel_word_count(chunks))
Map runs in parallel across workers. Reduce (merge) runs on the main process — typically fast, since each map result is already aggregated. The shape is identical to MapReduce/Spark; the only thing missing is the cluster.
13. A Benchmark Worth Internalising
The canonical demo: a CPU-bound function, run serially vs through a thread pool vs through a process pool.
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor import time def heavy(n): return sum(i * i for i in range(n)) WORK = [10_000_000] * 8 def serial(): return [heavy(n) for n in WORK] def threads(): with ThreadPoolExecutor(max_workers=8) as pool: return list(pool.map(heavy, WORK)) def processes(): with ProcessPoolExecutor(max_workers=8) as pool: return list(pool.map(heavy, WORK)) if __name__ == "__main__": for label, fn in [("serial", serial), ("threads", threads), ("processes", processes)]: t0 = time.perf_counter() fn() print(f"{label:>10}: {time.perf_counter()-t0:.2f}s")
On a typical 8-core laptop:
serial: ~6.0s threads: ~6.0s (no faster — GIL serialises the bytecode loop) processes: ~0.9s (~6.7× speedup — bounded by core count and pickling overhead)
Threads gave you zero. Processes gave you near-linear scaling. That gap is the GIL. If your workload looks like this, multiprocessing is the only standard-library answer.
(If heavy were numpy.sum(numpy.arange(n)**2) instead — the loop lives in C and releases the GIL — threads would scale too. Knowing which path your code takes is half the optimisation work.)
Common Mistakes
1. Threading CPU-bound code and expecting speedup
The most common misconception. Pure-Python loops do not parallelise with threads, full stop. Profile first; if the hot loop is Python bytecode, switch to processes or move the loop into NumPy/C.
2. Forgetting if __name__ == "__main__": on Windows/macOS
Endless subprocess spawning, CPU pegged, task manager full of python.exe. The guard is mandatory for any module that calls Process, Pool, or ProcessPoolExecutor.
3. Mutating "shared" state in subprocesses
counter = 0 def bump(): global counter; counter += 1 with ProcessPoolExecutor() as p: list(p.map(bump, range(10))) print(counter) # 0 — each child has its own `counter`
setup added so this can run · defines ProcessPoolExecutor
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def ProcessPoolExecutor(*_a, **_kw): print('-> ProcessPoolExecutor() called') return _AutoMock('ProcessPoolExecutor()')
Each subprocess gets a copy of the parent's state at spawn time. Mutations don't propagate back. Use Manager, Queue, or shared_memory — or design out shared state by returning values from workers and aggregating in the parent.
4. Deadlocks from inconsistent lock ordering
Thread 1 holds A, wants B. Thread 2 holds B, wants A. Both wait forever. The fix is discipline: pick a global order (e.g. by id(lock)) and always acquire in that order. Better: design out the need for two locks at once.
5. Pickle errors on closures, lambdas, local functions
def outer(): def inner(x): return x * 2 with ProcessPoolExecutor() as p: list(p.map(inner, range(10))) # PicklingError: can't pickle local function
setup added so this can run · defines ProcessPoolExecutor
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def ProcessPoolExecutor(*_a, **_kw): print('-> ProcessPoolExecutor() called') return _AutoMock('ProcessPoolExecutor()')
Process pools pickle the function to ship it to workers. Pickling needs the function to be importable at module top level. Move inner out, or use multiprocess (third-party fork using dill) — but module-level functions are the right answer.
6. Pool overhead dwarfing the work
Submitting 100,000 tiny tasks to a process pool spends more time pickling than computing. Two fixes: batch the work (each task does 1000 items instead of 1) or pass chunksize=N to pool.map so the pool batches for you.
7. Not joining/closing the pool
Always with ThreadPoolExecutor(...) as pool:. The with block calls shutdown(wait=True) on exit — workers drain, threads stop. Without it, workers can outlive the main thread and prevent the program from exiting.
🎯 Your Turn — parallel_map(fn, items, mode="thread")
Write a function that parallelises a map across either threads or processes, picked by a mode argument. Then benchmark it on a CPU-bound function and prove to yourself that only process mode actually speeds things up.
Requirements:
- Signature:
parallel_map(fn, items, mode="thread", max_workers=None). mode="thread"usesThreadPoolExecutor;mode="process"usesProcessPoolExecutor; anything else raisesValueError.- Returns results in input order (same as
pool.map). - If
max_workersisNone, let the executor choose (it defaults to a reasonable value). - Include a benchmark block (under
if __name__ == "__main__":) that times a CPU-bound function in serial, thread, and process mode.
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor def parallel_map(fn, items, mode="thread", max_workers=None): # TODO 1: pick the executor class based on `mode` # TODO 2: open the executor in a `with` block # TODO 3: return list(executor.map(fn, items)) ... def cpu_heavy(n): return sum(i * i for i in range(n)) if __name__ == "__main__": work = [2_000_000] * 8 # TODO 4: time serial baseline # TODO 5: time parallel_map(cpu_heavy, work, mode="thread") # TODO 6: time parallel_map(cpu_heavy, work, mode="process") # Print the three timings.
Hint 1 — Mapping mode to executor
A simpledict works: {"thread": ThreadPoolExecutor, "process": ProcessPoolExecutor}[mode]. Use .get(mode) and raise ValueError if the result is None. Or use a plain if/elif/else chain.
Hint 2 — Why the benchmark needs the __main__ guard
ProcessPoolExecutor spawns subprocesses that re-import your script. Without the guard, those subprocesses would also try to run the benchmark, which would spawn more subprocesses, which would... your machine catches fire. Every multiprocessing example you write must live under if __name__ == "__main__":.
Show full solution
from concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor import time def parallel_map(fn, items, mode="thread", max_workers=None): """Parallelise `fn` over `items` using threads or processes. Results returned in input order, like the built-in `map`. """ executor_cls = { "thread": ThreadPoolExecutor, "process": ProcessPoolExecutor, }.get(mode) if executor_cls is None: raise ValueError(f"mode must be 'thread' or 'process', got {mode!r}") with executor_cls(max_workers=max_workers) as pool: return list(pool.map(fn, items)) def cpu_heavy(n): """Pure-Python CPU loop — does NOT release the GIL.""" return sum(i * i for i in range(n)) def timed(label, fn): t0 = time.perf_counter() fn() elapsed = time.perf_counter() - t0 print(f"{label:>14}: {elapsed:.2f}s") if __name__ == "__main__": work = [2_000_000] * 8 timed("serial", lambda: [cpu_heavy(n) for n in work]) timed("thread pool", lambda: parallel_map(cpu_heavy, work, mode="thread")) timed("process pool", lambda: parallel_map(cpu_heavy, work, mode="process"))
Typical output on an 8-core machine:
serial: 1.20s thread pool: 1.22s <- no speedup; the GIL serialises bytecode process pool: 0.18s <- ~6.5x speedup; each process owns its own GIL
What this benchmark proves, in one screenful:
- Threads are useless here. Eight workers, one core's worth of throughput. The GIL forces the eight Python interpreters to take turns on the bytecode loop.
- Processes scale. Each worker is a separate interpreter with its own GIL, running on its own core. Speedup is sub-linear (pickling argument and result has overhead, and the OS doesn't give you all 8 cores at 100%) but the order of magnitude is right.
- The API is identical. Same
parallel_mapfunction, samepool.map, samewithblock. The one-wordmodechange is all it takes to switch concurrency models.
A more interesting variant: replace cpu_heavy with numpy.sum(numpy.arange(n) ** 2) — that loop runs in C and releases the GIL. Threads will then scale just as well as processes, and with no pickling overhead. The general rule: GIL-releasing work parallelises with threads; pure-Python work needs processes.
For real workloads, profile first. The right knob to turn is often "vectorise with NumPy" or "rewrite the hot loop in Cython," not "throw more workers at it." Concurrency is the answer when the work is already efficient and you simply need more of it at once.
What You Learned
- The GIL — one thread of Python bytecode at a time per process. Threads help with I/O and GIL-releasing C code; they don't help with pure-Python CPU work.
- Decision matrix: I/O + async libs → asyncio; I/O + blocking libs → threads; CPU-bound → processes.
threading.Thread,Lock,RLock,Semaphore— the low-level primitives. Always acquire underwith.queue.Queue— thread-safe producer/consumer. Almost always better than rolling your own with locks.threading.local()— per-thread state.contextvarsis the async-friendly modern alternative.concurrent.futures.ThreadPoolExecutor— the modern preferred thread API.submitreturns aFuture;mappreserves order;as_completedyields in completion order.concurrent.futures.ProcessPoolExecutor— same shape, separate processes, true parallelism. Arguments and results are pickled.multiprocessing.Pool— older API;imap_unorderedis still useful. New code: preferProcessPoolExecutor.- Sharing across processes:
Queue(messages),Manager(proxied state — slow),shared_memory(zero-copy bytes — fast, dangerous). if __name__ == "__main__":is mandatory with multiprocessing on Windows/macOS (spawn) — otherwise child processes re-import the script and respawn forever.- Map-reduce —
pool.map(map_fn, chunks)+ a serialreduce. The shape of every batch job you'll ever write. - Pickling tax — function arguments and return values cross process boundaries serialised. Batch small tasks (
chunksize=) and avoid large per-call payloads. - Pool overhead can exceed benefit for tiny tasks. Measure before parallelising.
Next: Type Hints in Depth — Protocol, TypeVar, Generic, Self, ParamSpec, and how the modern Python type system actually catches bugs at the boundary.