Grouping is everywhere: orders by customer, readings by sensor, files by extension. The hand-written version spends half its lines asking "have I seen this key before?". defaultdict answers that for you.
Before
orders = [("Ada", 30), ("Linus", 12), ("Ada", 55), ("Grace", 8), ("Linus", 20)] by_customer = {} for name, total in orders: if name not in by_customer: by_customer[name] = [] by_customer[name].append(total) print(by_customer)
{'Ada': [30, 55], 'Linus': [12, 20], 'Grace': [8]}After
from collections import defaultdict orders = [("Ada", 30), ("Linus", 12), ("Ada", 55), ("Grace", 8), ("Linus", 20)] by_customer = defaultdict(list) for name, total in orders: by_customer[name].append(total) print(dict(by_customer)) print({name: sum(t) for name, t in by_customer.items()})
{'Ada': [30, 55], 'Linus': [12, 20], 'Grace': [8]}
{'Ada': 85, 'Linus': 32, 'Grace': 8}Why it works
defaultdict(list) stores a factory, the list type itself. When you read a missing key, it calls the factory, stores the new empty list under that key, and hands it back, so .append always has something to append to. Any zero-argument callable works: defaultdict(int) for counting, defaultdict(set) for unique values per key.
from collections import defaultdict visits = [("home", "Ada"), ("docs", "Linus"), ("home", "Ada"), ("home", "Grace")] unique_visitors = defaultdict(set) for page, user in visits: unique_visitors[page].add(user) print({page: sorted(users) for page, users in unique_visitors.items()})
{'home': ['Ada', 'Grace'], 'docs': ['Linus']}When not to use it
Reading a missing key creates it. A lookup like if by_customer["Bob"]: quietly adds an empty list for Bob; use "Bob" in by_customer or .get("Bob") to check without creating. If you're only counting, Counter is the better fit, and if the data already comes sorted by key, itertools.groupby groups it without building the dict at all.