PythonMastery

Group rows with defaultdict(list)

defaultdict(list) creates the empty list the first time a key is seen, so grouping rows by city, customer or date is two lines with no key check.

Grouping is everywhere: orders by customer, readings by sensor, files by extension. The hand-written version spends half its lines asking "have I seen this key before?". defaultdict answers that for you.

Before

python
orders = [("Ada", 30), ("Linus", 12), ("Ada", 55), ("Grace", 8), ("Linus", 20)]

by_customer = {}
for name, total in orders:
    if name not in by_customer:
        by_customer[name] = []
    by_customer[name].append(total)

print(by_customer)
output
{'Ada': [30, 55], 'Linus': [12, 20], 'Grace': [8]}

After

python
from collections import defaultdict

orders = [("Ada", 30), ("Linus", 12), ("Ada", 55), ("Grace", 8), ("Linus", 20)]

by_customer = defaultdict(list)
for name, total in orders:
    by_customer[name].append(total)

print(dict(by_customer))
print({name: sum(t) for name, t in by_customer.items()})
output
{'Ada': [30, 55], 'Linus': [12, 20], 'Grace': [8]}
{'Ada': 85, 'Linus': 32, 'Grace': 8}

Why it works

defaultdict(list) stores a factory, the list type itself. When you read a missing key, it calls the factory, stores the new empty list under that key, and hands it back, so .append always has something to append to. Any zero-argument callable works: defaultdict(int) for counting, defaultdict(set) for unique values per key.

python
from collections import defaultdict

visits = [("home", "Ada"), ("docs", "Linus"), ("home", "Ada"), ("home", "Grace")]
unique_visitors = defaultdict(set)
for page, user in visits:
    unique_visitors[page].add(user)

print({page: sorted(users) for page, users in unique_visitors.items()})
output
{'home': ['Ada', 'Grace'], 'docs': ['Linus']}

When not to use it

Reading a missing key creates it. A lookup like if by_customer["Bob"]: quietly adds an empty list for Bob; use "Bob" in by_customer or .get("Bob") to check without creating. If you're only counting, Counter is the better fit, and if the data already comes sorted by key, itertools.groupby groups it without building the dict at all.

Learn it properly: collections: The Stdlib's Hidden Power, Dictionaries