Comprehensions: Lists, Dicts, Sets, Generators
1 · The lesson
readA comprehension is a one-line, declarative way to build a list, set, dict, or generator from another iterable. You describe what you want — the shape of each element and the source — and Python builds it. The imperative for loop describes how; the comprehension describes what.
Once you internalise the four forms, half the loops you used to write disappear.
1. The Four Forms
Same syntax, four containers:
nums = [1, 2, 3, 4, 5] squares_list = [n * n for n in nums] # list — [1, 4, 9, 16, 25] squares_set = {n * n for n in nums} # set — {1, 4, 9, 16, 25} squares_dict = {n: n * n for n in nums} # dict — {1: 1, 2: 4, ...} squares_gen = (n * n for n in nums) # generator — lazy iterator
The brackets pick the type. The rest is identical: expression for variable in iterable. The generator form doesn't build anything until you iterate it — more on that in Section 8.
2. Why Comprehensions
Compare the imperative form to the declarative one:
# Imperative — 4 lines, mutates an accumulator doubled = [] for n in nums: doubled.append(n * 2) # Declarative — one line, no accumulator, no .append doubled = [n * 2 for n in nums]
setup added so this can run · defines nums
nums = ["alpha", "beta", "gamma"]
The comprehension is shorter, allocates the list at known size (slightly faster), and reads as a single thought: "doubled values, one for each n in nums". You stop thinking about loop mechanics and start thinking in transformations.
3. Filtering with if
Append an if after the for to skip items:
evens = [n for n in range(20) if n % 2 == 0] # [0, 2, 4, 6, 8, 10, 12, 14, 16, 18] long_names = [name for name in ["Al", "Alice", "Bo", "Charlie"] if len(name) > 3] # ['Alice', 'Charlie']
You can chain multiple ifs — they're combined with logical AND:
divisible_by_six = [n for n in range(50) if n % 2 == 0 if n % 3 == 0] # same as divisible_by_six = [n for n in range(50) if n % 2 == 0 and n % 3 == 0]
Two ifs reads fine. Three is a code smell — break it up.
4. Conditional Expression Inside the Result
There are two places if can appear in a comprehension, and they mean different things.
nums = [-3, -1, 0, 2, 5] # Filtering — skips items positives = [n for n in nums if n > 0] # [2, 5] # Conditional expression — keeps every item, transforms based on a condition clamped = [n if n > 0 else 0 for n in nums] # [0, 0, 0, 2, 5]
The first drops negatives. The second replaces them with zero. The position of the if is what changes:
[expr for x in it if cond]—ifafter thefor→ filter.[expr_a if cond else expr_b for x in it]— ternary insideexpr→ transform.
You can combine them — filter and transform:
# Square positives only, drop the rest [n * n for n in nums if n > 0] # [4, 25] # Square positives, zero out negatives, keep all [n * n if n > 0 else 0 for n in nums] # [0, 0, 0, 4, 25]
setup added so this can run · defines nums
nums = [3, -1, 4, -1, 5]
5. Nested Comprehensions — Matrices
A comprehension whose expression is itself a comprehension builds a 2D structure.
# 3x4 multiplication table table = [[r * c for c in range(1, 5)] for r in range(1, 4)] # [[1, 2, 3, 4], # [2, 4, 6, 8], # [3, 6, 9, 12]]
Read it outside-in: for each row r, build a list of r * c for each column c. The outer for runs slowest; the inner runs fastest.
6. Flattening — Two fors, Same Line
When both fors sit at the top level, you get flattening, not nesting:
matrix = [[1, 2, 3], [4, 5, 6], [7, 8, 9]] flat = [item for row in matrix for item in row] # [1, 2, 3, 4, 5, 6, 7, 8, 9]
Read left-to-right exactly as you'd write the nested loop:
# Equivalent imperative form flat = [] for row in matrix: # first for for item in row: # second for flat.append(item) # the expression on the left
setup added so this can run · defines matrix
matrix = [[1, 2, 3], [4, 5, 6], [7, 8, 9]]
The order is the same in both forms — the comprehension just inverts where the body goes.
7. Set and Dict Comprehensions
Set comprehensions deduplicate as they build:
words = ["Hello", "world", "hello", "WORLD"] unique_lower = {w.lower() for w in words} # {'hello', 'world'}
Two operations — transform and dedup — in a single expression. The imperative form needs a loop and an explicit set.add call.
Dict comprehensions need a key and a value separated by ::
prices = {"apple": 1.20, "banana": 0.40, "cherry": 3.00}
# Transform values
shouty = {k: v for k, v in prices.items()} # straight copy
upper = {k.upper(): v for k, v in prices.items()} # {'APPLE': 1.2, ...}
# Swap keys and values
inverse = {v: k for k, v in prices.items()} # {1.2: 'apple', ...}
# Filter
cheap = {k: v for k, v in prices.items() if v < 1.0} # {'banana': 0.4}
# Build from scratch
squares = {n: n * n for n in range(5)} # {0: 0, 1: 1, 2: 4, 3: 9, 4: 16}The inverse trick ({v: k for k, v in d.items()}) is the canonical Pythonic way to flip a dict — provided your values are unique and hashable. Duplicate values silently collapse: the last one wins.
8. Generator Expressions — Lazy by Default
Swap the [ ] for ( ) and you get a generator — an iterator that yields one value at a time, computing on demand. It doesn't build the full sequence in memory.
# List — builds all 1,000,000 squares in memory total = sum([n * n for n in range(1_000_000)]) # Generator — yields one square at a time, sum consumes them as they come total = sum(n * n for n in range(1_000_000))
Same answer, vastly different memory profile. When you're feeding into sum, any, all, min, max, "".join(...), or any function that just iterates once, the generator form is almost always the right choice — and you can drop the outer parentheses:
# All of these work with no extra parens sum(n * n for n in range(100)) any(x < 0 for x in nums) all(s.isalpha() for s in words) max(len(line) for line in open("data.txt")) "\n".join(line.strip() for line in lines)
setup added so this can run · defines nums, words, lines
nums = [3, -1, 4, -1, 5] words = ["alpha", "beta", "gamma"] lines = ["alpha", "beta", "gamma"]
Generators are single-use. After you've iterated one, it's spent — you can't rewind. If you need to walk the data twice, use a list.
9. Walrus in a Comprehension (3.8+)
The := operator binds and returns in one step. Inside a comprehension, it lets you reuse an expensive computation without computing it twice — once for the filter, once for the result.
import re lines = ["age=42", "broken", "name=Alice", "also-broken"] pairs = [(m.group(1), m.group(2)) for line in lines if (m := re.match(r"(\w+)=(\w+)", line))] # [('age', '42'), ('name', 'Alice')]
setup added so this can run · defines m
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) m = _AutoMock('m')
Without the walrus you'd either call re.match twice or precompute with a regular loop. Use it for genuinely expensive or non-trivial expressions — for simple ones, the readability cost isn't worth it.
10. When Not to Use a Comprehension
Comprehensions are sharp tools. Reach for a regular for loop when:
- The logic spans more than one readable line. If your comprehension wraps onto three lines and you squint at it, the loop is clearer.
- You need multiple levels of
if/elifin the result expression. Ternary chains nested inside a comprehension are write-only code. - The body has side effects — file writes,
print, mutations. Comprehensions are for producing a collection. Throwing away a list ofNones just to run side effects is a misuse:
python
# BAD — builds and discards a list of Nones
[print(x) for x in items]
# Correct
for x in items:
print(x)
- You need exception handling per element. A
try/exceptdoesn't fit inside a comprehension.
The rule: a comprehension should describe one transformation in one line. The moment it grows, fall back to the loop.
Common Mistakes
1. Confusing filter-if with conditional-expression-if
nums = [-2, -1, 0, 1, 2] # Position matters — these do different things [n if n > 0 else 0 for n in nums] # [0, 0, 0, 1, 2] — transform [n for n in nums if n > 0] # [1, 2] — filter
Filter goes after the for. Conditional expression goes before the for, inside the result slot.
2. Using [ ] when you wanted ( )
# Materialises 10 million numbers in memory just to sum them total = sum([n for n in range(10_000_000)]) # Streams — constant memory total = sum(n for n in range(10_000_000))
If the comprehension is feeding into a single-pass consumer (sum, any, all, max, join...), prefer the generator form.
3. The readability cliff — nesting too deep
# Don't do this result = [[x * y for x in row if x > 0] for row in matrix if sum(row) > 10]
setup added so this can run · defines matrix, y
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) matrix = ["alpha", "beta", "gamma"] y = _AutoMock('y')
That's two fors, two ifs, and an arithmetic expression in one line. Six months from now, neither you nor anyone else will parse it. Break it into a function with a regular nested loop and a clear name.
4. Name shadowing — mostly fine, with one quirk
In Python 3, comprehensions have their own scope — the loop variable doesn't leak out:
i = "outer" squares = [i * i for i in range(5)] print(i) # 'outer' — not 4
Good — no surprise mutations of the outer i. But the iterable expression itself is evaluated in the enclosing scope:
i = [1, 2, 3] result = [x for x in i for x in range(x)] # the outer `i` is the source list
Tricky and unrecommended — just don't reuse names across the source and the loop variable.
5. Comprehensions for side effects
# Wrong — builds a useless list of Nones [file.write(line) for line in lines] # Right for line in lines: file.write(line)
setup added so this can run · defines lines, file
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) lines = ["alpha", "beta", "gamma"] file = _AutoMock('file')
If you're not going to use the result, don't build a result. Lists aren't free.
🎯 Your Turn — Pivot Table
Write pivot_table(rows, key, value) that groups a list of dicts by one field and collects another field's values into a list.
people = [
{"dept": "eng", "salary": 90},
{"dept": "sales", "salary": 60},
{"dept": "eng", "salary": 110},
{"dept": "sales", "salary": 80},
{"dept": "eng", "salary": 95},
]
pivot_table(people, key="dept", value="salary")
# {'eng': [90, 110, 95], 'sales': [60, 80]} setup added so this can run · defines pivot_table
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def pivot_table(*_a, **_kw): print('-> pivot_table() called') return _AutoMock('pivot_table()')
Skeleton:
def pivot_table(rows, key, value): # TODO 1: for each row, append row[value] under the bucket row[key] # TODO 2: return the resulting dict ...
Hint 1 — Grouping pattern
You can't write this purely as a dict comprehension — you need to append to existing lists as you go. Either build an empty dict and usedict.setdefault(k, []).append(v) inside a regular for, or import collections.defaultdict(list) and skip the setdefault dance.
Hint 2 — setdefault vs defaultdict
d.setdefault("eng", []).append(90) returns the existing list for "eng" (creating an empty one if missing), then appends to it. defaultdict(list) does the same thing automatically the first time you index a missing key — slightly cleaner if you'll do this pattern repeatedly.
Show full solution
Version A — setdefault, no imports:
def pivot_table(rows, key, value): result = {} for row in rows: result.setdefault(row[key], []).append(row[value]) return result
Version B — defaultdict, slightly cleaner:
from collections import defaultdict def pivot_table(rows, key, value): result = defaultdict(list) for row in rows: result[row[key]].append(row[value]) return dict(result) # cast back to plain dict for the caller
Version C — pure comprehension by way of a unique-keys pass:
def pivot_table(rows, key, value): return {k: [r[value] for r in rows if r[key] == k] for k in {r[key] for r in rows}}
Version C is the most "comprehension-y" but it's O(n × k) — for every unique key, it walks the whole row list again. Versions A and B are O(n). On 10 rows it doesn't matter; on 10 million it does. This is exactly the trade-off comprehensions force you to weigh: clever vs clear vs fast.
Test it:
people = [
{"dept": "eng", "salary": 90},
{"dept": "sales", "salary": 60},
{"dept": "eng", "salary": 110},
]
print(pivot_table(people, "dept", "salary"))
# {'eng': [90, 110], 'sales': [60]} setup added so this can run · defines pivot_table
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def pivot_table(*_a, **_kw): print('-> pivot_table() called') return _AutoMock('pivot_table()')
What You Learned
- Four forms — list
[ ], set{ }, dict{k: v}, generator( )— same syntax, different containers. - Filter
ifgoes after thefor; conditional-expressionifgoes before theforinside the result slot. - Nested comprehensions build matrices; flat double-
forflattens them. - Dict comprehensions transform, invert, or rebuild dicts in one line.
- Generator expressions are lazy — use them whenever a single-pass consumer (
sum,any,all,join...) is on the other end. - Stop reaching for a comprehension once it stops fitting on one readable line. The goal is clarity, not cleverness.
- Don't use comprehensions for side effects. Use a
forloop.
Next: Decorators — functions that wrap other functions to add cross-cutting behaviour.
Practice this
on practicepython.inShort exercises that run in your browser and tell you what your code actually did, not just whether a test passed.