PythonMastery
beginner 18 min read · lesson 14 of 19 in Python Fundamentals

Dictionaries

1 · The lesson

read

A dictionary maps keys to values. Where a list answers "what's at position 3?", a dict answers "what's the value for 'username'?" — and answers it in roughly constant time no matter how big the dict gets.

If lists are Python's workhorse, dicts are its nervous system. JSON, function keyword arguments, object attributes, configuration, caches — all dicts under the hood.


1. What a Dict Is

python
user = {
    "name": "Alice",
    "age": 28,
    "languages": ["Python", "TypeScript"],
}

print(user["name"])             # 'Alice'
print(user["languages"])        # ['Python', 'TypeScript']

Three properties to lock in:

  • Key-based — you look things up by key, not position.
  • Hash-backed — lookup, insertion, and deletion are O(1) on average.
  • Insertion-ordered — since Python 3.7, iteration yields keys in the order they were added.

Keys must be hashable — strings, numbers, tuples-of-hashables, booleans. Lists and dicts can't be keys.


2. Creating Dicts

python
a = {}                                          # empty dict
b = {"name": "Ada", "age": 36}                  # literal — most common
c = dict(name="Linus", age=54)                  # keyword form — keys must be valid identifiers
d = dict([("x", 1), ("y", 2)])                  # from a list of pairs
e = dict(zip(["a", "b", "c"], [1, 2, 3]))       # zip two parallel lists

print(b, c, d, e)

Note: {} is an empty dict, not an empty set. The empty set is set(). We'll get to Sets shortly.


3. Reading Values

Two ways. The choice matters.

python
user = {"name": "Alice", "age": 28}

print(user["name"])             # 'Alice'
# print(user["email"])          # KeyError: 'email'

print(user.get("email"))        # None         — never raises
print(user.get("email", "n/a")) # 'n/a'        — default if missing

d[key] says "I expect this key to be here; crash if it isn't." That's often the right behaviour — a missing key is a bug, and you want to know.

d.get(key) says "give me this if you have it, otherwise give me nothing." Use it when missing is normal.

Don't reflexively reach for .get() to suppress errors. A KeyError is information.


4. Setting & Updating

python
user = {"name": "Alice"}

user["age"] = 28                # add a key
user["name"] = "Alice T"        # overwrite a key
print(user)                     # {'name': 'Alice T', 'age': 28}

user.update({"age": 29, "city": "Hyderabad"})
print(user)                     # {'name': 'Alice T', 'age': 29, 'city': 'Hyderabad'}

update() merges another dict (or iterable of pairs) into this one — overwriting on key collision.


5. Removing Keys

python
d = {"a": 1, "b": 2, "c": 3}

del d["a"]                      # raises KeyError if missing

value = d.pop("b")              # value == 2; d is now {'c': 3}
missing = d.pop("z", None)      # returns None instead of raising

d.clear()                       # {}

pop is the one to reach for when you both want the value and want it gone.


6. Iterating

python
scores = {"Ada": 95, "Linus": 78, "Grace": 99}

for key in scores:              # iterating a dict yields its keys
    print(key)

for key in scores.keys():       # explicit — same result
    print(key)

for value in scores.values():
    print(value)

for name, score in scores.items():
    print(f"{name} scored {score}")

.items() is the one you'll use most. Unpacking the pair right in the for is the idiomatic shape.

.keys(), .values(), .items() return views — live windows onto the dict. They update if the dict updates, and they're cheap (no copy).


7. Membership

python
user = {"name": "Alice", "age": 28}

print("name" in user)           # True   — checks keys, not values
print("Alice" in user)          # False
print("Alice" in user.values()) # True   — be explicit if you mean values

x in d always checks keys. That's the fast, hash-backed operation. Searching values is linear — fine occasionally, not in a hot loop.


8. setdefault — Get or Create

The pattern "give me the list for this key, creating it if needed" is so common it has a method.

python
# Group words by first letter
words = ["apple", "ant", "banana", "berry", "cherry"]
groups = {}

for word in words:
    groups.setdefault(word[0], []).append(word)

print(groups)
# {'a': ['apple', 'ant'], 'b': ['banana', 'berry'], 'c': ['cherry']}

setdefault(key, default) returns the existing value if key is there, or inserts default and returns that. Either way you get back something to work with.

If you find yourself reaching for setdefault constantly, jump to defaultdict below — it's cleaner.


9. Counter and defaultdict

Two collections classes that earn their keep immediately.

Counter — count things.

python
from collections import Counter

text = "the quick brown fox jumps over the lazy dog the"
counts = Counter(text.split())

print(counts)                   # Counter({'the': 3, 'quick': 1, 'brown': 1, ...})
print(counts["the"])            # 3
print(counts["missing"])        # 0   — never raises, returns 0
print(counts.most_common(2))    # [('the', 3), ('quick', 1)]

Counter is a dict subclass — everything you know about dicts still works.

defaultdict — auto-create missing values.

python
from collections import defaultdict

groups = defaultdict(list)      # missing key → empty list, automatically
for word in ["apple", "ant", "banana", "berry"]:
    groups[word[0]].append(word)

print(dict(groups))             # {'a': ['apple', 'ant'], 'b': ['banana', 'berry']}

You pass the factory (the callable, not the value): list, int, set, or your own function. defaultdict(int) is the canonical counter when you want an exact dict shape rather than a Counter.


10. Dict Comprehensions

Same idea as list comprehensions, but you produce key-value pairs.

python
nums = [1, 2, 3, 4]
squares = {n: n * n for n in nums}
print(squares)                  # {1: 1, 2: 4, 3: 9, 4: 16}

# Invert a dict (assuming unique values)
original = {"a": 1, "b": 2, "c": 3}
inverted = {v: k for k, v in original.items()}
print(inverted)                 # {1: 'a', 2: 'b', 3: 'c'}

# Filter
prices = {"apple": 1.0, "steak": 12.0, "bread": 2.5}
cheap  = {k: v for k, v in prices.items() if v < 5}
print(cheap)                    # {'apple': 1.0, 'bread': 2.5}

11. Merging Dicts

python
defaults = {"theme": "dark", "font_size": 14}
overrides = {"font_size": 16, "language": "en"}

merged = {**defaults, **overrides}      # works on all modern Pythons
print(merged)                           # {'theme': 'dark', 'font_size': 16, 'language': 'en'}

merged2 = defaults | overrides          # Python 3.9+  — cleaner
print(merged2)                          # same result

On key collision, the right-hand dict wins. That's why overrides going second is correct for a settings pattern.

The | operator is the modern preference — it reads like "defaults, with overrides applied".


12. The Keys-Only Cousin — Sets

A set is essentially a dict with no values — just unique, hashable members.

python
seen = {"apple", "banana", "apple"}
print(seen)                     # {'apple', 'banana'}
print("apple" in seen)          # True — O(1)

Reach for a set when you need fast membership testing or uniqueness with no associated value. Full coverage in Sets.


Common Mistakes

1. KeyError vs .get() — using the wrong tool.

python
config = {"host": "localhost"}

port = config["port"]           # KeyError — and that's often what you want!
port = config.get("port", 5432) # silent default — only use when missing is genuinely OK

Suppressing a KeyError with .get() for required config is how you end up debugging a None three function calls away. Let it raise where the problem actually is.

2. Iterating .keys() when you wanted .items().

python
scores = {"Ada": 95, "Linus": 78}

for name in scores:
    print(name, scores[name])   # works, but you're paying for a second lookup

for name, score in scores.items():
    print(name, score)          # cleaner and faster

3. Mutating a dict while iterating it.

python
d = {"a": 1, "b": 2, "c": 3}
for key in d:
    if d[key] == 2:
        del d[key]              # RuntimeError: dictionary changed size during iteration

Iterate over a list of the keys: for key in list(d): — that snapshot is independent of the dict.

4. Using a list as a key.

python
# d = {[1, 2]: "oops"}          # TypeError: unhashable type: 'list'
d = {(1, 2): "fine"}            # tuples work — they're hashable

Anything mutable can't be a key. Use a tuple of immutables instead.

5. Assuming dict order doesn't matter.

Since Python 3.7, dicts preserve insertion order. Code that relies on this is fine on modern Python — but be aware that two dicts with the same keys and values in different insertion order still compare equal (== ignores order). Order matters for iteration; it does not matter for equality.


🎯 Your Turn — Build group_by

Write group_by(items, key_fn) that takes a list and a function. It calls key_fn(item) on each item to produce a group key, and returns a dict mapping each key to the list of items that produced it — order preserved.

python
group_by([1, 2, 3, 4, 5, 6], lambda n: n % 2)
# {1: [1, 3, 5], 0: [2, 4, 6]}

group_by(["apple", "ant", "berry", "banana", "cherry"], lambda s: s[0])
# {'a': ['apple', 'ant'], 'b': ['berry', 'banana'], 'c': ['cherry']}

group_by([], lambda x: x)
# {}
+ setup added so this can run · defines group_by
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def group_by(*_a, **_kw):
    print('-> group_by() called')
    return _AutoMock('group_by()')

It's the SQL GROUP BY for Python lists — and the building block behind a hundred real-world transformations (group orders by customer, log lines by level, students by grade).

Skeleton:

python
def group_by(items, key_fn):
    # TODO 1: start with an empty container
    # TODO 2: for each item, compute its key, append the item to that key's list
    # TODO 3: return a plain dict
    ...

print(group_by([1, 2, 3, 4, 5, 6], lambda n: n % 2))
Hint 1 — The "get or create the list" pattern You need "give me the list for this key, creating an empty one if it doesn't exist, then append". That's exactly what setdefault does — or, even cleaner, what defaultdict(list) was built for.
Hint 2 — Why convert at the end? A defaultdict behaves like a dict in almost every way, but printing one (and comparing in tests) shows the defaultdict(<class 'list'>, ...) wrapper. Wrapping the result in dict(...) before returning gives callers a plain dict — friendlier output and no surprise auto-creation downstream.
Show full solution
python
from collections import defaultdict

def group_by(items, key_fn):
    groups = defaultdict(list)
    for item in items:
        groups[key_fn(item)].append(item)
    return dict(groups)

# Sanity checks
print(group_by([1, 2, 3, 4, 5, 6], lambda n: n % 2))
# {1: [1, 3, 5], 0: [2, 4, 6]}

print(group_by(["apple", "ant", "berry", "banana", "cherry"], lambda s: s[0]))
# {'a': ['apple', 'ant'], 'b': ['berry', 'banana'], 'c': ['cherry']}

print(group_by([], lambda x: x))
# {}

Four lines. defaultdict(list) removes the "does this key exist yet?" boilerplate; the append just works. Converting to a plain dict at the end keeps the return type predictable for callers.

The setdefault version is one line longer and equally fine:

python
def group_by(items, key_fn):
    groups = {}
    for item in items:
        groups.setdefault(key_fn(item), []).append(item)
    return groups

Either is idiomatic. Pick defaultdict when you'll grow the groups across multiple loops; pick setdefault for one-shot grouping where importing feels heavy.


What You Learned

  • A dict is a hash-backed, insertion-ordered, key→value mapping. Lookup, insert, and delete are O(1) average.
  • Keys must be hashable — strings, numbers, tuples-of-hashables. Lists and dicts cannot be keys.
  • d[key] raises on missing; d.get(key, default) doesn't. Choose based on whether missing is a bug or a normal case.
  • Iterate with .items() for key-value pairs, .keys() or .values() when you only need one side.
  • setdefault and collections.defaultdict handle the get-or-create pattern. Counter handles counting.
  • Dict comprehensions: {k: f(v) for k, v in d.items() if cond}.
  • Merge with {**a, **b} or a | b (Python 3.9+) — right-hand side wins on collision.

Next: Tuples — the immutable cousin you've been hearing about — and then Sets for fast uniqueness checks.

Practice this

on practicepython.in

Short exercises that run in your browser and tell you what your code actually did, not just whether a test passed.