PythonMastery
intermediate 15 min read · lesson 4 of 13 in Python Intermediate

Comprehensions: Lists, Dicts, Sets, Generators

1 · The lesson

read

A comprehension is a one-line, declarative way to build a list, set, dict, or generator from another iterable. You describe what you want — the shape of each element and the source — and Python builds it. The imperative for loop describes how; the comprehension describes what.

Once you internalise the four forms, half the loops you used to write disappear.


1. The Four Forms

Same syntax, four containers:

python
nums = [1, 2, 3, 4, 5]

squares_list  = [n * n for n in nums]           # list        — [1, 4, 9, 16, 25]
squares_set   = {n * n for n in nums}           # set         — {1, 4, 9, 16, 25}
squares_dict  = {n: n * n for n in nums}        # dict        — {1: 1, 2: 4, ...}
squares_gen   = (n * n for n in nums)           # generator   — lazy iterator

The brackets pick the type. The rest is identical: expression for variable in iterable. The generator form doesn't build anything until you iterate it — more on that in Section 8.


2. Why Comprehensions

Compare the imperative form to the declarative one:

python
# Imperative — 4 lines, mutates an accumulator
doubled = []
for n in nums:
    doubled.append(n * 2)

# Declarative — one line, no accumulator, no .append
doubled = [n * 2 for n in nums]
+ setup added so this can run · defines nums
nums = ["alpha", "beta", "gamma"]

The comprehension is shorter, allocates the list at known size (slightly faster), and reads as a single thought: "doubled values, one for each n in nums". You stop thinking about loop mechanics and start thinking in transformations.


3. Filtering with if

Append an if after the for to skip items:

python
evens = [n for n in range(20) if n % 2 == 0]
#       [0, 2, 4, 6, 8, 10, 12, 14, 16, 18]

long_names = [name for name in ["Al", "Alice", "Bo", "Charlie"] if len(name) > 3]
#            ['Alice', 'Charlie']

You can chain multiple ifs — they're combined with logical AND:

python
divisible_by_six = [n for n in range(50) if n % 2 == 0 if n % 3 == 0]
# same as
divisible_by_six = [n for n in range(50) if n % 2 == 0 and n % 3 == 0]

Two ifs reads fine. Three is a code smell — break it up.


4. Conditional Expression Inside the Result

There are two places if can appear in a comprehension, and they mean different things.

python
nums = [-3, -1, 0, 2, 5]

# Filtering — skips items
positives = [n for n in nums if n > 0]                  # [2, 5]

# Conditional expression — keeps every item, transforms based on a condition
clamped   = [n if n > 0 else 0 for n in nums]           # [0, 0, 0, 2, 5]

The first drops negatives. The second replaces them with zero. The position of the if is what changes:

  • [expr for x in it if cond] — if after the for → filter.
  • [expr_a if cond else expr_b for x in it] — ternary inside expr → transform.

You can combine them — filter and transform:

python
# Square positives only, drop the rest
[n * n for n in nums if n > 0]                          # [4, 25]

# Square positives, zero out negatives, keep all
[n * n if n > 0 else 0 for n in nums]                   # [0, 0, 0, 4, 25]
+ setup added so this can run · defines nums
nums = [3, -1, 4, -1, 5]

5. Nested Comprehensions — Matrices

A comprehension whose expression is itself a comprehension builds a 2D structure.

python
# 3x4 multiplication table
table = [[r * c for c in range(1, 5)] for r in range(1, 4)]
# [[1, 2, 3, 4],
#  [2, 4, 6, 8],
#  [3, 6, 9, 12]]

Read it outside-in: for each row r, build a list of r * c for each column c. The outer for runs slowest; the inner runs fastest.


6. Flattening — Two fors, Same Line

When both fors sit at the top level, you get flattening, not nesting:

python
matrix = [[1, 2, 3], [4, 5, 6], [7, 8, 9]]

flat = [item for row in matrix for item in row]
# [1, 2, 3, 4, 5, 6, 7, 8, 9]

Read left-to-right exactly as you'd write the nested loop:

python
# Equivalent imperative form
flat = []
for row in matrix:                  # first for
    for item in row:                # second for
        flat.append(item)           # the expression on the left
+ setup added so this can run · defines matrix
matrix = [[1, 2, 3], [4, 5, 6], [7, 8, 9]]

The order is the same in both forms — the comprehension just inverts where the body goes.


7. Set and Dict Comprehensions

Set comprehensions deduplicate as they build:

python
words = ["Hello", "world", "hello", "WORLD"]
unique_lower = {w.lower() for w in words}
# {'hello', 'world'}

Two operations — transform and dedup — in a single expression. The imperative form needs a loop and an explicit set.add call.

Dict comprehensions need a key and a value separated by ::

python
prices = {"apple": 1.20, "banana": 0.40, "cherry": 3.00}

# Transform values
shouty = {k: v for k, v in prices.items()}              # straight copy
upper  = {k.upper(): v for k, v in prices.items()}      # {'APPLE': 1.2, ...}

# Swap keys and values
inverse = {v: k for k, v in prices.items()}             # {1.2: 'apple', ...}

# Filter
cheap = {k: v for k, v in prices.items() if v < 1.0}    # {'banana': 0.4}

# Build from scratch
squares = {n: n * n for n in range(5)}                  # {0: 0, 1: 1, 2: 4, 3: 9, 4: 16}

The inverse trick ({v: k for k, v in d.items()}) is the canonical Pythonic way to flip a dict — provided your values are unique and hashable. Duplicate values silently collapse: the last one wins.


8. Generator Expressions — Lazy by Default

Swap the [ ] for ( ) and you get a generator — an iterator that yields one value at a time, computing on demand. It doesn't build the full sequence in memory.

python
# List — builds all 1,000,000 squares in memory
total = sum([n * n for n in range(1_000_000)])

# Generator — yields one square at a time, sum consumes them as they come
total = sum(n * n for n in range(1_000_000))

Same answer, vastly different memory profile. When you're feeding into sum, any, all, min, max, "".join(...), or any function that just iterates once, the generator form is almost always the right choice — and you can drop the outer parentheses:

python
# All of these work with no extra parens
sum(n * n for n in range(100))
any(x < 0 for x in nums)
all(s.isalpha() for s in words)
max(len(line) for line in open("data.txt"))
"\n".join(line.strip() for line in lines)
+ setup added so this can run · defines nums, words, lines
nums = [3, -1, 4, -1, 5]
words = ["alpha", "beta", "gamma"]
lines = ["alpha", "beta", "gamma"]

Generators are single-use. After you've iterated one, it's spent — you can't rewind. If you need to walk the data twice, use a list.


9. Walrus in a Comprehension (3.8+)

The := operator binds and returns in one step. Inside a comprehension, it lets you reuse an expensive computation without computing it twice — once for the filter, once for the result.

python
import re

lines = ["age=42", "broken", "name=Alice", "also-broken"]
pairs = [(m.group(1), m.group(2))
         for line in lines
         if (m := re.match(r"(\w+)=(\w+)", line))]
# [('age', '42'), ('name', 'Alice')]
+ setup added so this can run · defines m
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

m = _AutoMock('m')

Without the walrus you'd either call re.match twice or precompute with a regular loop. Use it for genuinely expensive or non-trivial expressions — for simple ones, the readability cost isn't worth it.


10. When Not to Use a Comprehension

Comprehensions are sharp tools. Reach for a regular for loop when:

  • The logic spans more than one readable line. If your comprehension wraps onto three lines and you squint at it, the loop is clearer.
  • You need multiple levels of if/elif in the result expression. Ternary chains nested inside a comprehension are write-only code.
  • The body has side effects — file writes, print, mutations. Comprehensions are for producing a collection. Throwing away a list of Nones just to run side effects is a misuse:

python # BAD — builds and discards a list of Nones [print(x) for x in items] # Correct for x in items: print(x)

  • You need exception handling per element. A try/except doesn't fit inside a comprehension.

The rule: a comprehension should describe one transformation in one line. The moment it grows, fall back to the loop.


Common Mistakes

1. Confusing filter-if with conditional-expression-if

python
nums = [-2, -1, 0, 1, 2]

# Position matters — these do different things
[n if n > 0 else 0 for n in nums]       # [0, 0, 0, 1, 2]  — transform
[n for n in nums if n > 0]              # [1, 2]            — filter

Filter goes after the for. Conditional expression goes before the for, inside the result slot.

2. Using [ ] when you wanted ( )

python
# Materialises 10 million numbers in memory just to sum them
total = sum([n for n in range(10_000_000)])

# Streams — constant memory
total = sum(n for n in range(10_000_000))

If the comprehension is feeding into a single-pass consumer (sum, any, all, max, join...), prefer the generator form.

3. The readability cliff — nesting too deep

python
# Don't do this
result = [[x * y for x in row if x > 0] for row in matrix if sum(row) > 10]
+ setup added so this can run · defines matrix, y
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

matrix = ["alpha", "beta", "gamma"]
y = _AutoMock('y')

That's two fors, two ifs, and an arithmetic expression in one line. Six months from now, neither you nor anyone else will parse it. Break it into a function with a regular nested loop and a clear name.

4. Name shadowing — mostly fine, with one quirk

In Python 3, comprehensions have their own scope — the loop variable doesn't leak out:

python
i = "outer"
squares = [i * i for i in range(5)]
print(i)                                # 'outer' — not 4

Good — no surprise mutations of the outer i. But the iterable expression itself is evaluated in the enclosing scope:

python
i = [1, 2, 3]
result = [x for x in i for x in range(x)]   # the outer `i` is the source list

Tricky and unrecommended — just don't reuse names across the source and the loop variable.

5. Comprehensions for side effects

python
# Wrong — builds a useless list of Nones
[file.write(line) for line in lines]

# Right
for line in lines:
    file.write(line)
+ setup added so this can run · defines lines, file
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

lines = ["alpha", "beta", "gamma"]
file = _AutoMock('file')

If you're not going to use the result, don't build a result. Lists aren't free.


🎯 Your Turn — Pivot Table

Write pivot_table(rows, key, value) that groups a list of dicts by one field and collects another field's values into a list.

python
people = [
    {"dept": "eng",   "salary": 90},
    {"dept": "sales", "salary": 60},
    {"dept": "eng",   "salary": 110},
    {"dept": "sales", "salary": 80},
    {"dept": "eng",   "salary": 95},
]

pivot_table(people, key="dept", value="salary")
# {'eng': [90, 110, 95], 'sales': [60, 80]}
+ setup added so this can run · defines pivot_table
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def pivot_table(*_a, **_kw):
    print('-> pivot_table() called')
    return _AutoMock('pivot_table()')

Skeleton:

python
def pivot_table(rows, key, value):
    # TODO 1: for each row, append row[value] under the bucket row[key]
    # TODO 2: return the resulting dict
    ...
Hint 1 — Grouping pattern You can't write this purely as a dict comprehension — you need to append to existing lists as you go. Either build an empty dict and use dict.setdefault(k, []).append(v) inside a regular for, or import collections.defaultdict(list) and skip the setdefault dance.
Hint 2 — setdefault vs defaultdict d.setdefault("eng", []).append(90) returns the existing list for "eng" (creating an empty one if missing), then appends to it. defaultdict(list) does the same thing automatically the first time you index a missing key — slightly cleaner if you'll do this pattern repeatedly.
Show full solution

Version A — setdefault, no imports:

python
def pivot_table(rows, key, value):
    result = {}
    for row in rows:
        result.setdefault(row[key], []).append(row[value])
    return result

Version B — defaultdict, slightly cleaner:

python
from collections import defaultdict

def pivot_table(rows, key, value):
    result = defaultdict(list)
    for row in rows:
        result[row[key]].append(row[value])
    return dict(result)             # cast back to plain dict for the caller

Version C — pure comprehension by way of a unique-keys pass:

python
def pivot_table(rows, key, value):
    return {k: [r[value] for r in rows if r[key] == k]
            for k in {r[key] for r in rows}}

Version C is the most "comprehension-y" but it's O(n × k) — for every unique key, it walks the whole row list again. Versions A and B are O(n). On 10 rows it doesn't matter; on 10 million it does. This is exactly the trade-off comprehensions force you to weigh: clever vs clear vs fast.

Test it:

python
people = [
    {"dept": "eng",   "salary": 90},
    {"dept": "sales", "salary": 60},
    {"dept": "eng",   "salary": 110},
]
print(pivot_table(people, "dept", "salary"))
# {'eng': [90, 110], 'sales': [60]}
+ setup added so this can run · defines pivot_table
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def pivot_table(*_a, **_kw):
    print('-> pivot_table() called')
    return _AutoMock('pivot_table()')

What You Learned

  • Four forms — list [ ], set { }, dict {k: v}, generator ( ) — same syntax, different containers.
  • Filter if goes after the for; conditional-expression if goes before the for inside the result slot.
  • Nested comprehensions build matrices; flat double-for flattens them.
  • Dict comprehensions transform, invert, or rebuild dicts in one line.
  • Generator expressions are lazy — use them whenever a single-pass consumer (sum, any, all, join...) is on the other end.
  • Stop reaching for a comprehension once it stops fitting on one readable line. The goal is clarity, not cleverness.
  • Don't use comprehensions for side effects. Use a for loop.

Next: Decorators — functions that wrap other functions to add cross-cutting behaviour.

Practice this

on practicepython.in

Short exercises that run in your browser and tell you what your code actually did, not just whether a test passed.