Sets: Unordered, Unique, Lightning Fast
1 · The lesson
readA set is Python's collection type for "a bag of unique things, in no particular order". You met sets briefly in the Dictionaries lesson — this lesson gives them the focused treatment they deserve.
Sets are the right tool for three jobs: removing duplicates, checking if something exists, and comparing groups (intersections, unions, differences). For those jobs they're dramatically faster than lists.
1. Creating a Set
# Set literal — curly braces with items fruits = {"apple", "banana", "cherry"} print(fruits) # {'banana', 'cherry', 'apple'} — order varies # From an iterable unique_numbers = set([1, 2, 2, 3, 3, 3]) print(unique_numbers) # {1, 2, 3} # Empty set — must use set(), NOT {} empty_set = set() # {} would create an empty DICT print(type(empty_set)) # <class 'set'> print(type({})) # <class 'dict'>
That last gotcha is famous — {} looks like an empty set but is actually a dict. Always use set() for an empty set.
2. The Two Properties That Define a Set
Unique — duplicates are silently dropped:
votes = ["yes", "no", "yes", "yes", "no", "maybe"] unique_votes = set(votes) print(unique_votes) # {'yes', 'no', 'maybe'} print(len(unique_votes)) # 3
Unordered — items have no fixed position:
s = {"c", "a", "b"}
print(s) # order is unpredictable
# print(s[0]) # TypeError — sets are not indexableIf you need order, use a list. If you need uniqueness without caring about order, set is faster and cleaner.
3. Membership Checks — The Killer Feature
Checking x in collection is the operation sets are fastest at. For a list of 1 million items, in is ~500,000 times slower than for a set of the same size.
allowed_users = {"alice", "bob", "carol", "dave"}
# Lightning fast — O(1) regardless of set size
print("alice" in allowed_users) # True
print("eve" in allowed_users) # FalseThis is why sets are the right tool for "is this in the blocklist?" or "have I already seen this username?" patterns.
4. Adding and Removing Items
tags = {"python", "tutorial"}
tags.add("beginner") # add one item
print(tags) # {'python', 'tutorial', 'beginner'}
tags.update(["free", "ide"]) # add many items
print(tags)
tags.remove("tutorial") # remove — raises KeyError if missing
tags.discard("not_there") # remove — silent if missing
tags.pop() # remove and return an arbitrary item
# Empty the whole set
tags.clear()
print(tags) # set()discard is the forgiving version of remove. Use it when you don't care whether the item was there.
5. Set Math — The Other Killer Feature
Sets support real mathematical set operations. These are dramatically more readable than the equivalent list-comprehension versions.
team_a = {"alice", "bob", "carol", "dave"}
team_b = {"carol", "dave", "eve", "frank"}
# UNION — in either set
print(team_a | team_b)
# {'alice', 'bob', 'carol', 'dave', 'eve', 'frank'}
# INTERSECTION — in both sets
print(team_a & team_b)
# {'carol', 'dave'}
# DIFFERENCE — in A but not B
print(team_a - team_b)
# {'alice', 'bob'}
# SYMMETRIC DIFFERENCE — in either, but not both
print(team_a ^ team_b)
# {'alice', 'bob', 'eve', 'frank'}The same operations exist as named methods if you prefer (team_a.union(team_b), .intersection(), .difference(), .symmetric_difference()).
Subset / superset tests:
required = {"name", "email"}
provided = {"name", "email", "age"}
print(required.issubset(provided)) # True
print(provided.issuperset(required)) # True
print(required <= provided) # True — shorthand for issubset6. Real-World Patterns
Deduplicate a list (preserving... nothing about order):
emails = ["a@b.com", "c@d.com", "a@b.com", "e@f.com"] unique_emails = list(set(emails)) print(unique_emails) # ['a@b.com', 'c@d.com', 'e@f.com']
If you need to deduplicate AND keep the original order, use dict.fromkeys() instead (dicts preserve insertion order in Python 3.7+):
ordered_unique = list(dict.fromkeys(emails)) print(ordered_unique) # original order preserved
setup added so this can run · defines emails
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) emails = _AutoMock('emails')
Find who's new vs returning:
yesterday = {"alice", "bob", "carol"}
today = {"bob", "carol", "dave", "eve"}
new_users = today - yesterday
print(new_users) # {'dave', 'eve'}
returning_users = today & yesterday
print(returning_users) # {'bob', 'carol'}
churned = yesterday - today
print(churned) # {'alice'}That's three useful business metrics in three lines.
Find duplicates in a list:
items = ["a", "b", "c", "a", "d", "b"] seen = set() duplicates = set() for x in items: if x in seen: duplicates.add(x) seen.add(x) print(duplicates) # {'a', 'b'}
7. frozenset — The Immutable Cousin
A normal set is mutable. frozenset is the read-only version — same operations, but you can't add or remove. The benefit: it's hashable, so it can be a dictionary key or a member of another set.
flavors = frozenset(["vanilla", "chocolate"]) flavors_with_strawberry = frozenset(["vanilla", "chocolate", "strawberry"]) # Use frozensets as dict keys — a normal set CAN'T do this order_count = { flavors: 5, flavors_with_strawberry: 2, } print(order_count[flavors]) # 5
You won't reach for frozenset often, but when you need a hashable set, it's there.
8. When to Use What
| Need | Use |
|---|---|
| Ordered items, possibly repeating, items might change | List |
| Ordered items, fixed once created, used as a "record" | Tuple |
| Key → value lookup | Dictionary |
| Unique items, fast membership, set math | Set |
| Unique items, immutable, hashable | frozenset |
9. Mistakes You'll Hit
1. {} is a dict, not a set
x = {} # dict
y = set() # set2. Trying to put a mutable thing in a set
# {[1, 2], [3, 4]} # TypeError — lists aren't hashable {(1, 2), (3, 4)} # tuples are fine
3. Indexing a set
# s = {"a", "b"} # s[0] # TypeError — sets aren't ordered
If you need indexing, convert to a list first:
list(s)[0] (but the "first" item is unpredictable).
Mini-Program — Vocab Overlap Between Two Texts
text_a = "the quick brown fox jumps over the lazy dog" text_b = "the lazy fox sleeps under a brown tree" words_a = set(text_a.split()) words_b = set(text_b.split()) shared = words_a & words_b only_a = words_a - words_b only_b = words_b - words_a print(f"Shared words: {sorted(shared)}") print(f"Unique to A: {sorted(only_a)}") print(f"Unique to B: {sorted(only_b)}") print(f"Vocab overlap: {len(shared) / len(words_a | words_b):.1%}")
That's set math doing real text analysis in 8 lines.
🎯 Your Turn — Who Dropped the Course?
You have two attendance registers: the students enrolled in week 1, and those
still enrolled in week 5. Answer three questions with set operations rather than
loops.
Write course_changes(week1, week5) returning a tuple of three sorted lists:
students who dropped, students who joined late, and students who stayed
throughout.
week1 = ["asha", "biju", "chitra", "deepak"] week5 = ["asha", "chitra", "elango"] dropped → ["biju", "deepak"] joined → ["elango"] stayed → ["asha", "chitra"]
Skeleton:
def course_changes(week1, week5): a = set(week1) b = set(week5) # TODO 1: dropped = in week1 but not week5 # TODO 2: joined = in week5 but not week1 # TODO 3: stayed = in both # TODO 4: return all three as sorted lists ... dropped, joined, stayed = course_changes( ["asha", "biju", "chitra", "deepak"], ["asha", "chitra", "elango"], ) print("dropped:", dropped) # ['biju', 'deepak'] print("joined :", joined) # ['elango'] print("stayed :", stayed) # ['asha', 'chitra']
Hint 1 — Three operators, three answers
a - b is everything in a that is not in b.
b - a is the reverse. a & b is what appears in
both. Each question in this exercise is exactly one of those.
Hint 2 — Sets have no order, so sort on the way out
Printing a set gives you an arbitrary order that can change between runs, which makes output impossible to test.sorted(some_set) returns a list in
a predictable order — do it once, at the boundary where the data leaves your
function.
Show full solution
def course_changes(week1, week5): a = set(week1) b = set(week5) return sorted(a - b), sorted(b - a), sorted(a & b) dropped, joined, stayed = course_changes( ["asha", "biju", "chitra", "deepak"], ["asha", "chitra", "elango"], ) print("dropped:", dropped) # ['biju', 'deepak'] print("joined :", joined) # ['elango'] print("stayed :", stayed) # ['asha', 'chitra']
Written with loops this would be three nested passes and about fifteen lines. As
set algebra it is one line, and it reads almost exactly like the question you
were asked — which is the real argument for sets. Note that the answer also
silently handles duplicates in the input registers, because building a set drops
them.
Recap
- A set is an unordered collection of unique items.
- Create with
{1, 2, 3}orset(...). Useset()for an empty set —{}is a dict. - Sets are the fastest way to check membership (
x in s) — O(1) lookups. - Set math operators:
|union,&intersection,-difference,^symmetric difference. - Use frozenset when you need an immutable, hashable set.
Sets are one of those tools that, once you know them, you'll reach for constantly. "Are these emails unique?" "What overlaps between these two groups?" "Have I seen this user before?" — set, set, set.
Source: adapted from Python official documentation Section 5.4 (Sets). PSF License.
Practice this
on practicepython.inShort exercises that run in your browser and tell you what your code actually did, not just whether a test passed.