PythonMastery
intermediate 18 min read · lesson 11 of 13 in Python Intermediate

Itertools: Composable Iterator Building Blocks

1 · The lesson

read

The itertools module is the Lego set of looping in Python. Each function returns a lazy iterator that consumes input one element at a time and yields one element at a time. They compose — the output of one is the input of the next — and they almost never build intermediate lists. That makes them memory-efficient on huge or infinite sequences, and lethal at one-line problems that would otherwise need nested loops and accumulators.

Once you know the dozen functions in this lesson, you'll catch yourself reaching for them weekly.


1. Why Itertools

Three properties shared by everything in the module:

1. Lazy — nothing is computed until you iterate. itertools.count() represents 0, 1, 2, 3, … forever in constant memory.
2. Composable — chain(a, b) returns an iterator; pass that to islice(..., 10) and you get the first ten of the concatenated stream, still without materialising either input.
3. Single-use — like all generators, once iterated they're exhausted. list(iter_obj) materialises if you need to keep the result.

python
import itertools as it

# Lazy — none of this is computed until consumed
stream = it.islice(it.count(start=1), 5)        # iterator over [1, 2, 3, 4, 5]
print(list(stream))                             # [1, 2, 3, 4, 5]
print(list(stream))                             # [] — already exhausted

That last detail catches everyone once. After iterating, the iterator is spent. Reuse the factory (it.count, it.islice) to get a fresh one, or materialise into a list.


2. The Infinite Ones — count, cycle, repeat

Three generators that never stop on their own. Always pair them with islice, takewhile, zip, or some other terminator.

python
import itertools as it

# count(start=0, step=1) — like range() with no stop
for i in it.count(start=10, step=2):
    if i > 20:
        break
    print(i)                                    # 10, 12, 14, 16, 18, 20

# cycle(iter) — loops forever
colours = it.cycle(["red", "green", "blue"])
for _ in range(7):
    print(next(colours))                        # red, green, blue, red, green, blue, red

# repeat(x, n=infinite) — same value, optionally bounded
list(it.repeat("hi", 3))                        # ['hi', 'hi', 'hi']

# Common pattern — number rows in a CSV
rows = ["a", "b", "c", "d"]
list(zip(it.count(1), rows))                    # [(1,'a'), (2,'b'), (3,'c'), (4,'d')]

zip(it.count(1), rows) is exactly enumerate(rows, start=1) — the canonical idiom. But count works in places enumerate doesn't, like combining with takewhile for "keep generating until a condition fails".


3. chain and chain.from_iterable — Concatenate Iterables

chain(*iterables) flattens one level by walking through each iterable in turn.

python
import itertools as it

list(it.chain([1, 2], [3, 4], (5, 6)))         # [1, 2, 3, 4, 5, 6]
list(it.chain("ab", "cd", "ef"))               # ['a', 'b', 'c', 'd', 'e', 'f']

chain.from_iterable(iter_of_iters) takes a single iterable of iterables — useful when you have a list of lists already:

python
matrix = [[1, 2, 3], [4, 5], [6, 7, 8, 9]]

list(it.chain.from_iterable(matrix))           # [1, 2, 3, 4, 5, 6, 7, 8, 9]

# Equivalent comprehension
[x for row in matrix for x in row]
+ setup added so this can run · defines it
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

it = _AutoMock('it')

For one or two levels of nesting, both forms work. For chaining many streams, from_iterable shines:

python
files = ["a.txt", "b.txt", "c.txt"]
all_lines = it.chain.from_iterable(open(f) for f in files)
for line in all_lines:                         # streams every line, never holds all files in memory
    process(line)
+ setup added so this can run · defines process, it
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def process(*_a, **_kw):
    print('-> process() called')
    return _AutoMock('process()')
it = _AutoMock('it')

4. takewhile and dropwhile — Split at the First Failure

takewhile(pred, iter) yields items while the predicate is true, then stops — even if later items would also pass.

python
import itertools as it

nums = [1, 3, 5, 4, 6, 7, 9]

list(it.takewhile(lambda n: n % 2 == 1, nums))  # [1, 3, 5]   — stops at 4
list(it.dropwhile(lambda n: n % 2 == 1, nums))  # [4, 6, 7, 9] — drops until first even, keeps rest

dropwhile is the mirror: it skips items while the predicate is true, then yields everything from the first failure onwards (predicate is not re-checked after the first miss).

Use case — strip leading comments from a file without reading it twice:

python
lines = ["# header", "# notes", "data1", "data2", "# inline comment"]
data = list(it.dropwhile(lambda l: l.startswith("#"), lines))
# ['data1', 'data2', '# inline comment']   — inline comment is kept!
+ setup added so this can run · defines it
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

it = _AutoMock('it')

This is not the same as filter. filter checks every element. takewhile/dropwhile short-circuit at the first transition.


5. islice — Slice Any Iterable

Lists support lst[2:10:2]. Generators don't. islice brings slice-like behaviour to any iterator:

python
import itertools as it

# islice(iter, stop)               — first N items
list(it.islice(it.count(), 5))                 # [0, 1, 2, 3, 4]

# islice(iter, start, stop)        — items in a range
list(it.islice(range(20), 5, 10))              # [5, 6, 7, 8, 9]

# islice(iter, start, stop, step)  — with stride
list(it.islice(range(20), 0, 20, 3))           # [0, 3, 6, 9, 12, 15, 18]

# stop=None means run to exhaustion
list(it.islice("abcdefgh", 2, None))           # ['c', 'd', 'e', 'f', 'g', 'h']

Two limits versus list slicing:

  • No negative indices — you can't ask for "the last five" of an iterator; it doesn't know its length.
  • It consumes the underlying iterator. After it.islice(stream, 5), the next item in stream is the sixth, not the first.

6. zip_longest — Zip Without Truncation

Plain zip stops at the shortest iterable. zip_longest runs until the longest is exhausted, filling missing values with fillvalue.

python
from itertools import zip_longest

names  = ["Alice", "Bob", "Charlie"]
scores = [90, 85]

list(zip(names, scores))                        # [('Alice', 90), ('Bob', 85)]   — Charlie dropped
list(zip_longest(names, scores, fillvalue=0))   # [('Alice', 90), ('Bob', 85), ('Charlie', 0)]

Use it when missing data is meaningful — survey responses, sparse columns, partial migrations. The fillvalue is a single value applied to all missing slots in all shorter iterables.


7. groupby — Group Consecutive Equal-Key Items

groupby(iter, key=fn) walks the iterable and yields (key, group_iterator) pairs whenever the key changes. It only merges adjacent runs. If you want to group globally, sort first.

python
from itertools import groupby

# Already sorted by first letter — groupby works as expected
words = ["apple", "ant", "banana", "berry", "cherry"]
for key, group in groupby(words, key=lambda w: w[0]):
    print(key, list(group))
# a ['apple', 'ant']
# b ['banana', 'berry']
# c ['cherry']

# NOT sorted — groupby silently fragments
words = ["apple", "banana", "ant", "berry", "cherry"]
for key, group in groupby(words, key=lambda w: w[0]):
    print(key, list(group))
# a ['apple']
# b ['banana']
# a ['ant']          — broken out as its own group!
# b ['berry']
# c ['cherry']

The fix is always the same — sort by the key first:

python
words.sort(key=lambda w: w[0])
groups = {k: list(g) for k, g in groupby(words, key=lambda w: w[0])}
# {'a': ['ant', 'apple'], 'b': ['banana', 'berry'], 'c': ['cherry']}
+ setup added so this can run · defines words, groupby
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

words = _AutoMock('words')
def groupby(*_a, **_kw):
    print('-> groupby() called')
    return _AutoMock('groupby()')

For pure group-and-collect with no need for the order of groups, collections.defaultdict(list) (one pass, no sort) is often simpler. groupby is for when consecutiveness is the semantic — runs of equal pixels, time-windowed events, log lines from the same request.

The other gotcha: the inner group is an iterator that shares state with the outer groupby. Consume it before moving on (list(group) in the example), or it'll be empty by the time you try.


8. accumulate — Running Totals (and More)

accumulate(iter, func=operator.add) yields the running result.

python
from itertools import accumulate
import operator

list(accumulate([1, 2, 3, 4]))                  # [1, 3, 6, 10]   — running sum (default)
list(accumulate([1, 2, 3, 4], operator.mul))    # [1, 2, 6, 24]   — running product
list(accumulate([3, 1, 4, 1, 5, 9, 2, 6], max)) # [3, 3, 4, 4, 5, 9, 9, 9]   — running max

# Bank balance — start with 100, apply daily changes
changes = [+50, -30, +100, -75]
list(accumulate(changes, initial=100))          # [100, 150, 120, 220, 145]  (3.8+)

The initial= keyword (Python 3.8+) prepends a starting value — useful for bank balances, prefix sums starting at 0, anything that needs a seed.

Gotcha: the function must be a binary operation (acc, x) -> new_acc, and order of arguments matters for non-commutative ops. accumulate([2, 1, 3], lambda a, b: a - b) gives [2, 1, -2], not [2, -1, -4].


9. product — Nested Loops as One Call

product(*iterables, repeat=1) is the Cartesian product — every combination of one item from each iterable.

python
from itertools import product

# Pairs of dice rolls
list(product([1, 2, 3], [1, 2, 3]))
# [(1,1), (1,2), (1,3), (2,1), (2,2), (2,3), (3,1), (3,2), (3,3)]

# Equivalent nested-loop form
[(a, b) for a in [1, 2, 3] for b in [1, 2, 3]]

# repeat — Cartesian product of the same iterable with itself
list(product("ab", repeat=3))
# [('a','a','a'), ('a','a','b'), ('a','b','a'), ..., ('b','b','b')]   — all 3-letter ab-strings

# Three-level loop over a grid
for x, y, z in product(range(3), range(3), range(3)):
    ...     # 27 iterations, no nesting

product(repeat=n) is how you write "all possible passwords of length n from this alphabet" or "every coordinate in this n-dimensional grid" without an n-deep nested loop.


10. permutations and combinations

When order matters → permutations. When it doesn't → combinations.

python
from itertools import permutations, combinations

# All orderings of 3 letters, length 2
list(permutations("abc", 2))
# [('a','b'), ('a','c'), ('b','a'), ('b','c'), ('c','a'), ('c','b')]   — 6 = 3*2

# All subsets of size 2 — order doesn't matter, (a,b) == (b,a)
list(combinations("abc", 2))
# [('a','b'), ('a','c'), ('b','c')]   — 3 = C(3,2)

# combinations_with_replacement — same element can repeat
from itertools import combinations_with_replacement
list(combinations_with_replacement("abc", 2))
# [('a','a'), ('a','b'), ('a','c'), ('b','b'), ('b','c'), ('c','c')]

These are your weapons for combinatorial problems — subset-sum, anagram detection, brute-force searches over small spaces. Note the size grows fast: permutations(range(10), 10) is 3.6 million; for range(12) it's 479 million. Use islice to peek at a few without materialising.


11. pairwise (Python 3.10+) — Sliding Window of 2

A surprisingly common operation that used to require boilerplate.

python
from itertools import pairwise

list(pairwise([1, 2, 3, 4, 5]))                 # [(1,2), (2,3), (3,4), (4,5)]

# Detect price changes day-to-day
prices = [100, 102, 99, 105, 110]
deltas = [b - a for a, b in pairwise(prices)]   # [2, -3, 6, 5]

For windows wider than 2, no built-in exists — write a generator with collections.deque(maxlen=n), or use the itertools recipes from the docs.


Common Mistakes

1. Forgetting itertools returns iterators, not lists

python
import itertools as it

result = it.chain([1, 2], [3, 4])
print(result)                                   # <itertools.chain object at 0x...>
print(len(result))                              # TypeError — no len on iterators

# Fix — materialise if you need length, indexing, or reuse
result = list(it.chain([1, 2], [3, 4]))

Don't reflexively list() everything — that throws away the memory benefit. But the moment you need len, indexing, or to iterate twice, list it.

2. groupby without pre-sorting

python
from itertools import groupby

records = [
    ("Alice", "eng"), ("Bob", "sales"), ("Carol", "eng"),
]

# BUG — groups Alice and Carol separately because Bob sits between
{k: [name for name, _ in g] for k, g in groupby(records, key=lambda r: r[1])}
# {'eng': ['Alice'], 'sales': ['Bob'], 'eng': ['Carol']}  — last 'eng' wins, Alice lost

# Fix — sort by the same key first
records.sort(key=lambda r: r[1])
{k: [name for name, _ in g] for k, g in groupby(records, key=lambda r: r[1])}
# {'eng': ['Alice', 'Carol'], 'sales': ['Bob']}

Silent data loss is the worst kind of bug. If consecutive-grouping isn't what you want, reach for defaultdict(list) instead.

3. Iterating cycle() without a stop condition

python
import itertools as it

for c in it.cycle("abc"):
    print(c)                                    # never ends — Ctrl+C

Pair every infinite iterator with islice, takewhile, zip against a finite iterable, or an explicit break. Two consecutive infinite iterators (zip(cycle(a), cycle(b))) also runs forever — the shortest is still infinite.

4. accumulate with non-commutative ops

python
from itertools import accumulate

list(accumulate([10, 1, 2, 3], lambda a, b: a - b))
# [10, 9, 7, 4]    — 10, 10-1=9, 9-2=7, 7-3=4

Many people assume this gives [10, 1-10, 2-9, 3-7] or similar. It doesn't — accumulate passes (running_total, next_item) in that order. For subtraction, division, and string formatting, write a small test before trusting your mental model.

5. Trying to slice an iterator with [:n]

python
import itertools as it

stream = it.count()
first_five = stream[:5]                         # TypeError — iterators don't support slicing

# Use islice instead
first_five = list(it.islice(stream, 5))

[:n] is a list/tuple/string feature, not a general iterator feature. islice is the iterator-shaped slicer.


🎯 Your Turn — Chunked

Write chunked(iterable, n) — a generator that yields successive n-item lists from any iterable. The final chunk may be shorter than n.

python
list(chunked(range(10), 3))         # [[0, 1, 2], [3, 4, 5], [6, 7, 8], [9]]
list(chunked("abcdefgh", 2))        # [['a', 'b'], ['c', 'd'], ['e', 'f'], ['g', 'h']]
list(chunked([], 4))                # []
+ setup added so this can run · defines chunked
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def chunked(*_a, **_kw):
    print('-> chunked() called')
    return _AutoMock('chunked()')

Use itertools.islice over a shared iterator inside a loop — that's the canonical pattern.

Skeleton:

python
from itertools import islice

def chunked(iterable, n):
    # TODO 1: get an iterator over the input  (iter(iterable))
    # TODO 2: loop forever, pulling the next n items via islice
    # TODO 3: stop when islice returns an empty chunk
    # TODO 4: otherwise yield the chunk as a list
    ...
Hint 1 — Sharing one iterator across calls The trick is to call iter(iterable) once to get a single iterator, then call islice(iterator, n) repeatedly on the same iterator. Each islice consumes up to n items from where the last one left off.
Hint 2 — Knowing when to stop list(islice(it, n)) returns an empty list once it is exhausted. Use while True, materialise each chunk into a list, and break when the chunk is empty (or yield and break before yielding the empty).
Show full solution
python
from itertools import islice

def chunked(iterable, n):
    """Yield successive n-sized lists from `iterable`. Final chunk may be shorter."""
    if n < 1:
        raise ValueError("n must be at least 1")
    it = iter(iterable)
    while True:
        chunk = list(islice(it, n))
        if not chunk:
            return
        yield chunk


print(list(chunked(range(10), 3)))
# [[0, 1, 2], [3, 4, 5], [6, 7, 8], [9]]

print(list(chunked("abcdefgh", 2)))
# [['a', 'b'], ['c', 'd'], ['e', 'f'], ['g', 'h']]

print(list(chunked([], 4)))
# []

Why the single it = iter(iterable)? Because islice consumes items from whatever iterator you give it. If you wrote list(islice(iterable, n)) inside the loop, every call would start a fresh iterator over the same list — infinite loop on the first chunk.

A one-liner using iter's two-argument form:

python
from itertools import islice

def chunked(iterable, n):
    it = iter(iterable)
    return iter(lambda: list(islice(it, n)), [])

iter(callable, sentinel) calls callable() repeatedly and stops when it returns sentinel. Here, lambda: list(islice(it, n)) returns the next chunk, and [] (an empty list) signals exhaustion. Same behaviour, twice the cleverness — keep the loop version for shared code.

Bonus — chunked is now in the standard library as itertools.batched(iterable, n) from Python 3.12 onwards. Yields tuples instead of lists, otherwise the same. If you're on 3.12+:

python
from itertools import batched
list(batched("abcdefgh", 2))
# [('a', 'b'), ('c', 'd'), ('e', 'f'), ('g', 'h')]

Writing it yourself is still the better learning exercise — and useful on every Python before 3.12.


What You Learned

  • Itertools functions are lazy iterators — composable, memory-efficient, single-use.
  • The infinite trio: count, cycle, repeat — always pair with a stop condition.
  • chain concatenates iterables; chain.from_iterable flattens one level.
  • takewhile/dropwhile split at the first failure; filter checks every element.
  • islice brings slice-like behaviour to any iterator (no negative indices, consumes the source).
  • zip_longest(fillvalue=...) for uneven iterables.
  • groupby groups consecutive equal-key items — sort first if you want global groups.
  • accumulate for running totals/maxes/products; mind argument order for non-commutative ops.
  • product flattens n-deep loops; permutations/combinations for combinatorial searches.
  • pairwise (3.10+) for size-2 sliding windows.
  • Always check whether you need to list(...) the result; iterators are single-use.

Next: Decorators — applying the higher-order patterns from lambdas to wrap functions with cross-cutting behaviour.

Practice this

on practicepython.in

Short exercises that run in your browser and tell you what your code actually did, not just whether a test passed.