PythonMastery
intermediate 18 min read · lesson 6 of 13 in Python Intermediate

Generators: Lazy Iteration with yield

1 · The lesson

read

Building a 10-million-item list so you can iterate over it once is a waste of memory and a waste of time. Most of the time you don't need the whole list — you need the next item, then the one after that, then the one after that. A generator produces items on demand: compute the next one when asked, then pause until asked again.

yield is the keyword that turns a function into a generator. Once you've internalised it, you can stream a 50 GB log file through ten chained transformations using a few megabytes of RAM. The same primitive powers itertools, file iteration, async iteration, and every "pipeline" pattern in Python.


1. The Problem With Eager Lists

python
def squares(n):
    return [x*x for x in range(n)]      # builds the whole list in memory

total = sum(squares(10_000_000))         # ~400 MB before sum even starts

The list comprehension allocates every square, all at once, even though sum only needs one at a time. For large n you're either slow or out of memory — for no reason. The fix is laziness.

python
def squares(n):
    for x in range(n):
        yield x*x                        # produce ONE value, then pause

total = sum(squares(10_000_000))         # constant memory

yield is the entire change. The function now returns a generator object that produces 10 million squares on demand, one at a time, and forgets each one as soon as sum has consumed it.


2. yield Turns a Function Into a Generator

The presence of any yield in a function body makes the whole function a generator function. Calling it doesn't run the body — it returns a generator object that you iterate.

python
def count_up_to(n):
    print("starting")
    i = 1
    while i <= n:
        yield i
        i += 1
    print("done")

g = count_up_to(3)
print(g)                            # <generator object count_up_to at 0x...>
                                    # NOTHING printed yet — body hasn't run

print(next(g))                      # starting   |   1
print(next(g))                      #            |   2
print(next(g))                      #            |   3
print(next(g))                      # done       |   StopIteration

Each next(g) runs the body until the next yield, hands you the yielded value, and pauses — preserving every local variable, the instruction pointer, even the call stack. The next next(g) resumes from the exact line after the yield.

When the function falls off the end (or hits return), Python raises StopIteration to signal the generator is exhausted.


3. Iterating Generators

You almost never call next() directly. for, list(), sum(), min(), max(), any(), all() — and most of itertools — all consume iterables the same way: call next() in a loop, catch StopIteration, stop.

python
def count_up_to(n):
    i = 1
    while i <= n:
        yield i
        i += 1

for x in count_up_to(5):
    print(x)                        # 1, 2, 3, 4, 5

print(list(count_up_to(5)))         # [1, 2, 3, 4, 5]
print(sum(count_up_to(100)))        # 5050

A for loop is the canonical way to drain a generator. Reach for next() only when you want exactly one value (e.g. peek at the head, then loop the rest).


4. Generator Expressions

Same shape as a list comprehension, but with round brackets instead of square. Lazy by default.

python
squares_list = [x*x for x in range(1_000_000)]      # builds 1M-element list
squares_gen  = (x*x for x in range(1_000_000))      # builds a generator

print(sum(squares_list))            # works
print(sum(squares_gen))             # also works, constant memory

When a generator expression is the only argument to a function, you can drop the parentheses — the function call's own parens are enough:

python
total = sum(x*x for x in range(1_000_000))          # no double parens needed
maximum = max(len(line) for line in open("log.txt"))

Curly braces give you {x*x for x in range(5)} — a set comprehension. With key: value, it's a dict comprehension. Generator expressions are the only round-bracketed comprehension. Memorise the shape.


5. yield from — Delegating to a Sub-Iterator

When you want to yield every item from another iterable, the long form is a for loop:

python
def chain_two(a, b):
    for x in a:
        yield x
    for x in b:
        yield x

yield from collapses that into one line:

python
def chain_two(a, b):
    yield from a
    yield from b

print(list(chain_two([1, 2, 3], "ab")))     # [1, 2, 3, 'a', 'b']

yield from also propagates send(), throw(), and return values from sub-generators — useful when composing them. For now, treat it as the clean way to fan out one generator into another.


6. Memory — The Whole Reason

python
import sys

list_comp = [x*x for x in range(1_000_000)]
gen_exp   = (x*x for x in range(1_000_000))

print(sys.getsizeof(list_comp))     # 8_448_728   (~8.4 MB)
print(sys.getsizeof(gen_exp))       #       208   (208 bytes)

The generator is roughly 40,000× smaller. It holds the recipe — the loop state and the expression — not the results. Each result is computed, consumed, and discarded.

sys.getsizeof only measures the outer object, not what it contains. The list's 8 MB is the array of pointers to the int objects; the generator's 208 B is the generator frame. Either way, the asymmetry is dramatic.


7. Real Generators — Reading a Huge File

The classic Pythonic idiom. open() itself returns an iterator of lines — already lazy.

python
def read_lines(path):
    with open(path, encoding="utf-8") as f:
        for line in f:                       # one line at a time, not whole file
            yield line.rstrip("\n")

def only_errors(lines):
    for line in lines:
        if "ERROR" in line:
            yield line

def with_lineno(lines):
    for i, line in enumerate(lines, start=1):
        yield f"{i:6d}: {line}"

# Pipeline: read → filter → annotate → print
for entry in with_lineno(only_errors(read_lines("server.log"))):
    print(entry)

A 10 GB log? Same code. Same memory. Each line passes through every stage exactly once and is then garbage-collected. This pattern — generators feeding generators — is the UNIX pipe of Python.


8. Pipelines — Composing Generators

python
def numbers():
    n = 1
    while True:                              # infinite generator — fine, it's lazy
        yield n
        n += 1

def squared(source):
    for x in source:
        yield x * x

def take(source, n):
    for i, x in enumerate(source):
        if i >= n:
            return
        yield x

pipeline = take(squared(numbers()), 5)
print(list(pipeline))                        # [1, 4, 9, 16, 25]

Three stages — produce, transform, limit — composed by function call. Nothing is computed until list() starts pulling. Infinite generators are completely safe as long as something downstream terminates (take, itertools.islice, a break, etc.).

itertools is full of pipeline-ready generators: islice, chain, groupby, takewhile, dropwhile, tee, accumulate. See the Itertools lesson — most of it is generator composition.


9. send(), throw(), close() — The Coroutine Side

Generators were extended in PEP 342 to be two-way. The caller can send values into the generator at each yield:

python
def echo():
    while True:
        received = yield
        print(f"got: {received}")

g = echo()
next(g)                                      # prime the generator (run to first yield)
g.send("hello")                              # got: hello
g.send("world")                              # got: world
g.close()                                    # raise GeneratorExit inside, ends it

g.throw(SomeException) raises an exception inside the generator at the paused yield, letting it handle or propagate. g.close() raises GeneratorExit, giving the generator a chance to clean up.

This dual-direction yield was the original "coroutine" mechanism. In modern Python, async/await replaces nearly every use case. Recognise the syntax if you see it in older code; otherwise, don't reach for it.


10. itertools — Generators in the Standard Library

Most of itertools is generator functions implemented in C — fast, lazy, composable.

python
from itertools import islice, chain, count, takewhile

# First 5 squares of an infinite counter
print(list(islice((x*x for x in count(1)), 5)))    # [1, 4, 9, 16, 25]

# Concatenate without copying
print(list(chain([1, 2], (3, 4), {5, 6})))         # [1, 2, 3, 4, 5, 6]

# Take while a condition holds
print(list(takewhile(lambda x: x < 10, count(1)))) # [1, 2, 3, 4, 5, 6, 7, 8, 9]

When you find yourself writing a complicated while loop with manual state, check itertools first — there's often a one-liner. See the Itertools lesson.


11. Async Generators — One Sentence Ahead

async def plus yield gives you an async generator, consumed with async for. Used for streaming async I/O — paginated APIs, websockets, server-sent events. You'll meet them when you reach asyncio. The mental model is identical: lazy, one-at-a-time, paused between yields — just awaitable.


Common Mistakes

1. Treating a generator like a list — it's one-shot

python
g = (x*x for x in range(5))
print(list(g))                      # [0, 1, 4, 9, 16]
print(list(g))                      # []          — already exhausted!

A generator can be iterated exactly once. The second loop sees a drained iterator. If you need to iterate twice:

python
# Option A: keep the recipe and re-call it
def squares(n):
    for x in range(n):
        yield x*x

print(list(squares(5)))             # fresh
print(list(squares(5)))             # fresh again

# Option B: materialise once
data = list(x*x for x in range(5))  # now a real list, iterable any number of times

2. len(generator) doesn't work

python
g = (x for x in range(5))
print(len(g))                       # TypeError: object of type 'generator' has no len()

Generators don't know their length — they're producing on demand. If you need a count, either materialise (len(list(g)) — but that drains it) or count as you go (sum(1 for _ in g)).

3. Using a generator when you need to re-iterate

python
results = (process(item) for item in source)
if any(r > threshold for r in results):     # iterates once
    for r in results:                       # already empty!
        ...
+ setup added so this can run · defines process, source, threshold
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def process(*_a, **_kw):
    print('-> process() called')
    return _AutoMock('process()')
source = ["alpha", "beta", "gamma"]
threshold = _AutoMock('threshold')

If you'll re-iterate, materialise once (results = list(...)) or write a generator function you can call repeatedly. itertools.tee can split one generator into N independent iterators, but it buffers and is rarely the right answer.

4. Confusing comprehension brackets

python
[x for x in r]          # list
(x for x in r)          # generator
{x for x in r}          # set
{k: v for k, v in r}    # dict
+ setup added so this can run · defines r
r = [("alpha", 1), ("beta", 2), ("gamma", 3)]

Round brackets in a comprehension always mean generator expression, never a tuple. The "tuple comprehension" doesn't exist as a syntax — tuple(x for x in r) is how you build one.

5. Holding files open without with

python
def lines(path):
    f = open(path)
    for line in f:
        yield line
    f.close()                       # never reached if the consumer breaks early

If the consumer stops iterating partway (break, an exception, islice), f.close() never runs. Use a with block — GeneratorExit raised on g.close() will trigger the context manager and close the file.

python
def lines(path):
    with open(path) as f:           # always closes, even on early exit
        for line in f:
            yield line

🎯 Your Turn — Stream a CSV in Chunks

Write read_csv_chunks(path, chunk_size=1000) that yields lists of up to chunk_size rows from a CSV file. Each row is a list of strings. Pair it with csv.reader (no need to parse manually). The function must:

  • be a generator (use yield, not return),
  • never hold more than chunk_size + 1 rows in memory at once,
  • yield a final partial chunk if the file's row count isn't a multiple of chunk_size,
  • use a with block to close the file even if the consumer stops early,
  • skip the header row if has_header=True is passed.
python
import csv

def read_csv_chunks(path, chunk_size=1000, has_header=False):
    # TODO 1: open the file inside a `with` block
    # TODO 2: build a csv.reader over the file
    # TODO 3: if has_header, advance past the first row
    # TODO 4: accumulate rows into a chunk list; yield when it reaches chunk_size
    # TODO 5: after the loop, yield any remaining partial chunk
    ...

# Usage with itertools.islice to grab only the first few chunks:
from itertools import islice
for chunk in islice(read_csv_chunks("big.csv", chunk_size=500, has_header=True), 3):
    print(f"got {len(chunk)} rows; first row: {chunk[0]}")
Hint 1 — Skipping the header A csv.reader is itself an iterator. next(reader) consumes one row and discards it — exactly what you want for the header. Wrap in a conditional: if has_header: next(reader, None) (the None default avoids StopIteration on an empty file).
Hint 2 — Don't forget the leftover Inside the loop, append each row to chunk; when len(chunk) == chunk_size, yield chunk and reset to []. After the loop ends, chunk may still hold rows — yield it once more if non-empty, otherwise the last partial batch is silently dropped.
Show full solution
python
import csv

def read_csv_chunks(path, chunk_size=1000, has_header=False):
    """Yield lists of up to chunk_size rows from a CSV file."""
    with open(path, newline="", encoding="utf-8") as f:
        reader = csv.reader(f)
        if has_header:
            next(reader, None)                  # skip header if present

        chunk = []
        for row in reader:
            chunk.append(row)
            if len(chunk) == chunk_size:
                yield chunk
                chunk = []

        if chunk:                               # final partial chunk
            yield chunk


# Demo with a tiny in-memory CSV
import io, csv as _csv
sample = io.StringIO("name,age\nalice,30\nbob,25\ncarol,28\ndan,40\nellen,22\n")
# (use the generator pattern the same way; here we read from StringIO)

def read_csv_chunks_from(fileobj, chunk_size, has_header=False):
    reader = _csv.reader(fileobj)
    if has_header: next(reader, None)
    chunk = []
    for row in reader:
        chunk.append(row)
        if len(chunk) == chunk_size:
            yield chunk
            chunk = []
    if chunk:
        yield chunk

print(list(read_csv_chunks_from(sample, chunk_size=2, has_header=True)))
# [[['alice', '30'], ['bob', '25']],
#  [['carol', '28'], ['dan', '40']],
#  [['ellen', '22']]]

Key properties of this solution:

  • Memory bounded — at any moment, you hold one chunk (≤ chunk_size rows) plus whatever the consumer hasn't yet released. A 50 GB CSV with chunk_size=1000 runs in megabytes.
  • Safe early exit — the with block closes the file even if the consumer does break after one chunk. Compose with itertools.islice(read_csv_chunks(path), 3) to pull only the first three chunks; the file closes cleanly when the slice is exhausted.
  • Final partial chunk — the if chunk: yield chunk after the loop is the line beginners forget. Without it, the last N rows (where N < chunk_size) silently vanish.

This pattern — read_in_chunks → transform → write_in_chunks — is the foundation of every streaming ETL job, log processor, and "too big to fit in RAM" data pipeline you'll ever build. Pandas calls the same idea chunksize= in read_csv; Spark and Dask generalise it across machines. The primitive is the same generator you just wrote.


What You Learned

  • A function with yield is a generator function. Calling it returns a generator object; the body doesn't run until you iterate.
  • Each yield pauses the function, hands a value to the consumer, and resumes from the same line on the next next().
  • Generator expressions (x for x in ...) — same shape as a list comp, but lazy. Round brackets only.
  • yield from sub delegates to a sub-iterator in one line.
  • Generators use constant memory regardless of the size of the sequence they produce.
  • Compose generators into pipelines (take(squared(numbers()))) — the UNIX-pipe pattern.
  • Generators are one-shot — iterate once and they're done. Re-call the generator function, or materialise with list(...), if you need to iterate again.
  • len() doesn't work on generators. Neither does indexing. They're streams, not collections.
  • itertools is the standard library's generator toolkit. Use it before writing manual loops.
  • Always wrap file I/O inside a generator in with open(...) so cleanup runs on early exit.
  • send(), throw(), close() exist — rarely needed; async/await is the modern equivalent.

Next: Lambda Expressions — the one-line anonymous function, and where it earns its keep next to map, filter, and key=.

Practice this

on practicepython.in

Short exercises that run in your browser and tell you what your code actually did, not just whether a test passed.