Dictionaries
1 · The lesson
readA dictionary maps keys to values. Where a list answers "what's at position 3?", a dict answers "what's the value for 'username'?" — and answers it in roughly constant time no matter how big the dict gets.
If lists are Python's workhorse, dicts are its nervous system. JSON, function keyword arguments, object attributes, configuration, caches — all dicts under the hood.
1. What a Dict Is
user = {
"name": "Alice",
"age": 28,
"languages": ["Python", "TypeScript"],
}
print(user["name"]) # 'Alice'
print(user["languages"]) # ['Python', 'TypeScript']Three properties to lock in:
- Key-based — you look things up by key, not position.
- Hash-backed — lookup, insertion, and deletion are O(1) on average.
- Insertion-ordered — since Python 3.7, iteration yields keys in the order they were added.
Keys must be hashable — strings, numbers, tuples-of-hashables, booleans. Lists and dicts can't be keys.
2. Creating Dicts
a = {} # empty dict
b = {"name": "Ada", "age": 36} # literal — most common
c = dict(name="Linus", age=54) # keyword form — keys must be valid identifiers
d = dict([("x", 1), ("y", 2)]) # from a list of pairs
e = dict(zip(["a", "b", "c"], [1, 2, 3])) # zip two parallel lists
print(b, c, d, e)Note: {} is an empty dict, not an empty set. The empty set is set(). We'll get to Sets shortly.
3. Reading Values
Two ways. The choice matters.
user = {"name": "Alice", "age": 28}
print(user["name"]) # 'Alice'
# print(user["email"]) # KeyError: 'email'
print(user.get("email")) # None — never raises
print(user.get("email", "n/a")) # 'n/a' — default if missingd[key] says "I expect this key to be here; crash if it isn't." That's often the right behaviour — a missing key is a bug, and you want to know.
d.get(key) says "give me this if you have it, otherwise give me nothing." Use it when missing is normal.
Don't reflexively reach for .get() to suppress errors. A KeyError is information.
4. Setting & Updating
user = {"name": "Alice"}
user["age"] = 28 # add a key
user["name"] = "Alice T" # overwrite a key
print(user) # {'name': 'Alice T', 'age': 28}
user.update({"age": 29, "city": "Hyderabad"})
print(user) # {'name': 'Alice T', 'age': 29, 'city': 'Hyderabad'}update() merges another dict (or iterable of pairs) into this one — overwriting on key collision.
5. Removing Keys
d = {"a": 1, "b": 2, "c": 3}
del d["a"] # raises KeyError if missing
value = d.pop("b") # value == 2; d is now {'c': 3}
missing = d.pop("z", None) # returns None instead of raising
d.clear() # {}pop is the one to reach for when you both want the value and want it gone.
6. Iterating
scores = {"Ada": 95, "Linus": 78, "Grace": 99}
for key in scores: # iterating a dict yields its keys
print(key)
for key in scores.keys(): # explicit — same result
print(key)
for value in scores.values():
print(value)
for name, score in scores.items():
print(f"{name} scored {score}").items() is the one you'll use most. Unpacking the pair right in the for is the idiomatic shape.
.keys(), .values(), .items() return views — live windows onto the dict. They update if the dict updates, and they're cheap (no copy).
7. Membership
user = {"name": "Alice", "age": 28}
print("name" in user) # True — checks keys, not values
print("Alice" in user) # False
print("Alice" in user.values()) # True — be explicit if you mean valuesx in d always checks keys. That's the fast, hash-backed operation. Searching values is linear — fine occasionally, not in a hot loop.
8. setdefault — Get or Create
The pattern "give me the list for this key, creating it if needed" is so common it has a method.
# Group words by first letter words = ["apple", "ant", "banana", "berry", "cherry"] groups = {} for word in words: groups.setdefault(word[0], []).append(word) print(groups) # {'a': ['apple', 'ant'], 'b': ['banana', 'berry'], 'c': ['cherry']}
setdefault(key, default) returns the existing value if key is there, or inserts default and returns that. Either way you get back something to work with.
If you find yourself reaching for setdefault constantly, jump to defaultdict below — it's cleaner.
9. Counter and defaultdict
Two collections classes that earn their keep immediately.
Counter — count things.
from collections import Counter text = "the quick brown fox jumps over the lazy dog the" counts = Counter(text.split()) print(counts) # Counter({'the': 3, 'quick': 1, 'brown': 1, ...}) print(counts["the"]) # 3 print(counts["missing"]) # 0 — never raises, returns 0 print(counts.most_common(2)) # [('the', 3), ('quick', 1)]
Counter is a dict subclass — everything you know about dicts still works.
defaultdict — auto-create missing values.
from collections import defaultdict groups = defaultdict(list) # missing key → empty list, automatically for word in ["apple", "ant", "banana", "berry"]: groups[word[0]].append(word) print(dict(groups)) # {'a': ['apple', 'ant'], 'b': ['banana', 'berry']}
You pass the factory (the callable, not the value): list, int, set, or your own function. defaultdict(int) is the canonical counter when you want an exact dict shape rather than a Counter.
10. Dict Comprehensions
Same idea as list comprehensions, but you produce key-value pairs.
nums = [1, 2, 3, 4] squares = {n: n * n for n in nums} print(squares) # {1: 1, 2: 4, 3: 9, 4: 16} # Invert a dict (assuming unique values) original = {"a": 1, "b": 2, "c": 3} inverted = {v: k for k, v in original.items()} print(inverted) # {1: 'a', 2: 'b', 3: 'c'} # Filter prices = {"apple": 1.0, "steak": 12.0, "bread": 2.5} cheap = {k: v for k, v in prices.items() if v < 5} print(cheap) # {'apple': 1.0, 'bread': 2.5}
11. Merging Dicts
defaults = {"theme": "dark", "font_size": 14}
overrides = {"font_size": 16, "language": "en"}
merged = {**defaults, **overrides} # works on all modern Pythons
print(merged) # {'theme': 'dark', 'font_size': 16, 'language': 'en'}
merged2 = defaults | overrides # Python 3.9+ — cleaner
print(merged2) # same resultOn key collision, the right-hand dict wins. That's why overrides going second is correct for a settings pattern.
The | operator is the modern preference — it reads like "defaults, with overrides applied".
12. The Keys-Only Cousin — Sets
A set is essentially a dict with no values — just unique, hashable members.
seen = {"apple", "banana", "apple"}
print(seen) # {'apple', 'banana'}
print("apple" in seen) # True — O(1)Reach for a set when you need fast membership testing or uniqueness with no associated value. Full coverage in Sets.
Common Mistakes
1. KeyError vs .get() — using the wrong tool.
config = {"host": "localhost"}
port = config["port"] # KeyError — and that's often what you want!
port = config.get("port", 5432) # silent default — only use when missing is genuinely OKSuppressing a
KeyError with .get() for required config is how you end up debugging a None three function calls away. Let it raise where the problem actually is.
2. Iterating .keys() when you wanted .items().
scores = {"Ada": 95, "Linus": 78}
for name in scores:
print(name, scores[name]) # works, but you're paying for a second lookup
for name, score in scores.items():
print(name, score) # cleaner and faster3. Mutating a dict while iterating it.
d = {"a": 1, "b": 2, "c": 3}
for key in d:
if d[key] == 2:
del d[key] # RuntimeError: dictionary changed size during iterationIterate over a list of the keys:
for key in list(d): — that snapshot is independent of the dict.
4. Using a list as a key.
# d = {[1, 2]: "oops"} # TypeError: unhashable type: 'list' d = {(1, 2): "fine"} # tuples work — they're hashable
Anything mutable can't be a key. Use a tuple of immutables instead.
5. Assuming dict order doesn't matter.
Since Python 3.7, dicts preserve insertion order. Code that relies on this is fine on modern Python — but be aware that two dicts with the same keys and values in different insertion order still compare equal (== ignores order). Order matters for iteration; it does not matter for equality.
🎯 Your Turn — Build group_by
Write group_by(items, key_fn) that takes a list and a function. It calls key_fn(item) on each item to produce a group key, and returns a dict mapping each key to the list of items that produced it — order preserved.
group_by([1, 2, 3, 4, 5, 6], lambda n: n % 2) # {1: [1, 3, 5], 0: [2, 4, 6]} group_by(["apple", "ant", "berry", "banana", "cherry"], lambda s: s[0]) # {'a': ['apple', 'ant'], 'b': ['berry', 'banana'], 'c': ['cherry']} group_by([], lambda x: x) # {}
setup added so this can run · defines group_by
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def group_by(*_a, **_kw): print('-> group_by() called') return _AutoMock('group_by()')
It's the SQL GROUP BY for Python lists — and the building block behind a hundred real-world transformations (group orders by customer, log lines by level, students by grade).
Skeleton:
def group_by(items, key_fn): # TODO 1: start with an empty container # TODO 2: for each item, compute its key, append the item to that key's list # TODO 3: return a plain dict ... print(group_by([1, 2, 3, 4, 5, 6], lambda n: n % 2))
Hint 1 — The "get or create the list" pattern
You need "give me the list for this key, creating an empty one if it doesn't exist, then append". That's exactly whatsetdefault does — or, even cleaner, what defaultdict(list) was built for.
Hint 2 — Why convert at the end?
Adefaultdict behaves like a dict in almost every way, but printing one (and comparing in tests) shows the defaultdict(<class 'list'>, ...) wrapper. Wrapping the result in dict(...) before returning gives callers a plain dict — friendlier output and no surprise auto-creation downstream.
Show full solution
from collections import defaultdict def group_by(items, key_fn): groups = defaultdict(list) for item in items: groups[key_fn(item)].append(item) return dict(groups) # Sanity checks print(group_by([1, 2, 3, 4, 5, 6], lambda n: n % 2)) # {1: [1, 3, 5], 0: [2, 4, 6]} print(group_by(["apple", "ant", "berry", "banana", "cherry"], lambda s: s[0])) # {'a': ['apple', 'ant'], 'b': ['berry', 'banana'], 'c': ['cherry']} print(group_by([], lambda x: x)) # {}
Four lines. defaultdict(list) removes the "does this key exist yet?" boilerplate; the append just works. Converting to a plain dict at the end keeps the return type predictable for callers.
The setdefault version is one line longer and equally fine:
def group_by(items, key_fn): groups = {} for item in items: groups.setdefault(key_fn(item), []).append(item) return groups
Either is idiomatic. Pick defaultdict when you'll grow the groups across multiple loops; pick setdefault for one-shot grouping where importing feels heavy.
What You Learned
- A dict is a hash-backed, insertion-ordered, key→value mapping. Lookup, insert, and delete are O(1) average.
- Keys must be hashable — strings, numbers, tuples-of-hashables. Lists and dicts cannot be keys.
d[key]raises on missing;d.get(key, default)doesn't. Choose based on whether missing is a bug or a normal case.- Iterate with
.items()for key-value pairs,.keys()or.values()when you only need one side. setdefaultandcollections.defaultdicthandle the get-or-create pattern.Counterhandles counting.- Dict comprehensions:
{k: f(v) for k, v in d.items() if cond}. - Merge with
{**a, **b}ora | b(Python 3.9+) — right-hand side wins on collision.
Next: Tuples — the immutable cousin you've been hearing about — and then Sets for fast uniqueness checks.
Practice this
on practicepython.inShort exercises that run in your browser and tell you what your code actually did, not just whether a test passed.