PythonMastery
beginner 12 min read · lesson 16 of 19 in Python Fundamentals

Sets: Unordered, Unique, Lightning Fast

1 · The lesson

read

A set is Python's collection type for "a bag of unique things, in no particular order". You met sets briefly in the Dictionaries lesson — this lesson gives them the focused treatment they deserve.

Sets are the right tool for three jobs: removing duplicates, checking if something exists, and comparing groups (intersections, unions, differences). For those jobs they're dramatically faster than lists.


1. Creating a Set

python
# Set literal — curly braces with items
fruits = {"apple", "banana", "cherry"}
print(fruits)                       # {'banana', 'cherry', 'apple'} — order varies

# From an iterable
unique_numbers = set([1, 2, 2, 3, 3, 3])
print(unique_numbers)               # {1, 2, 3}

# Empty set — must use set(), NOT {}
empty_set = set()                   # {} would create an empty DICT
print(type(empty_set))              # <class 'set'>
print(type({}))                     # <class 'dict'>

That last gotcha is famous — {} looks like an empty set but is actually a dict. Always use set() for an empty set.


2. The Two Properties That Define a Set

Unique — duplicates are silently dropped:

python
votes = ["yes", "no", "yes", "yes", "no", "maybe"]
unique_votes = set(votes)
print(unique_votes)                 # {'yes', 'no', 'maybe'}
print(len(unique_votes))            # 3

Unordered — items have no fixed position:

python
s = {"c", "a", "b"}
print(s)                            # order is unpredictable
# print(s[0])                       # TypeError — sets are not indexable

If you need order, use a list. If you need uniqueness without caring about order, set is faster and cleaner.


3. Membership Checks — The Killer Feature

Checking x in collection is the operation sets are fastest at. For a list of 1 million items, in is ~500,000 times slower than for a set of the same size.

python
allowed_users = {"alice", "bob", "carol", "dave"}

# Lightning fast — O(1) regardless of set size
print("alice" in allowed_users)     # True
print("eve" in allowed_users)       # False

This is why sets are the right tool for "is this in the blocklist?" or "have I already seen this username?" patterns.


4. Adding and Removing Items

python
tags = {"python", "tutorial"}

tags.add("beginner")                # add one item
print(tags)                         # {'python', 'tutorial', 'beginner'}

tags.update(["free", "ide"])        # add many items
print(tags)

tags.remove("tutorial")             # remove — raises KeyError if missing
tags.discard("not_there")           # remove — silent if missing
tags.pop()                          # remove and return an arbitrary item

# Empty the whole set
tags.clear()
print(tags)                         # set()

discard is the forgiving version of remove. Use it when you don't care whether the item was there.


5. Set Math — The Other Killer Feature

Sets support real mathematical set operations. These are dramatically more readable than the equivalent list-comprehension versions.

python
team_a = {"alice", "bob", "carol", "dave"}
team_b = {"carol", "dave", "eve", "frank"}

# UNION — in either set
print(team_a | team_b)
# {'alice', 'bob', 'carol', 'dave', 'eve', 'frank'}

# INTERSECTION — in both sets
print(team_a & team_b)
# {'carol', 'dave'}

# DIFFERENCE — in A but not B
print(team_a - team_b)
# {'alice', 'bob'}

# SYMMETRIC DIFFERENCE — in either, but not both
print(team_a ^ team_b)
# {'alice', 'bob', 'eve', 'frank'}

The same operations exist as named methods if you prefer (team_a.union(team_b), .intersection(), .difference(), .symmetric_difference()).

Subset / superset tests:

python
required = {"name", "email"}
provided = {"name", "email", "age"}

print(required.issubset(provided))      # True
print(provided.issuperset(required))    # True
print(required <= provided)             # True — shorthand for issubset

6. Real-World Patterns

Deduplicate a list (preserving... nothing about order):

python
emails = ["a@b.com", "c@d.com", "a@b.com", "e@f.com"]
unique_emails = list(set(emails))
print(unique_emails)                # ['a@b.com', 'c@d.com', 'e@f.com']

If you need to deduplicate AND keep the original order, use dict.fromkeys() instead (dicts preserve insertion order in Python 3.7+):

python
ordered_unique = list(dict.fromkeys(emails))
print(ordered_unique)               # original order preserved
+ setup added so this can run · defines emails
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

emails = _AutoMock('emails')

Find who's new vs returning:

python
yesterday = {"alice", "bob", "carol"}
today = {"bob", "carol", "dave", "eve"}

new_users = today - yesterday
print(new_users)                    # {'dave', 'eve'}

returning_users = today & yesterday
print(returning_users)              # {'bob', 'carol'}

churned = yesterday - today
print(churned)                      # {'alice'}

That's three useful business metrics in three lines.

Find duplicates in a list:

python
items = ["a", "b", "c", "a", "d", "b"]
seen = set()
duplicates = set()
for x in items:
    if x in seen:
        duplicates.add(x)
    seen.add(x)
print(duplicates)                   # {'a', 'b'}

7. frozenset — The Immutable Cousin

A normal set is mutable. frozenset is the read-only version — same operations, but you can't add or remove. The benefit: it's hashable, so it can be a dictionary key or a member of another set.

python
flavors = frozenset(["vanilla", "chocolate"])
flavors_with_strawberry = frozenset(["vanilla", "chocolate", "strawberry"])

# Use frozensets as dict keys — a normal set CAN'T do this
order_count = {
    flavors: 5,
    flavors_with_strawberry: 2,
}
print(order_count[flavors])         # 5

You won't reach for frozenset often, but when you need a hashable set, it's there.


8. When to Use What

NeedUse
Ordered items, possibly repeating, items might changeList
Ordered items, fixed once created, used as a "record"Tuple
Key → value lookupDictionary
Unique items, fast membership, set mathSet
Unique items, immutable, hashablefrozenset

9. Mistakes You'll Hit

1. {} is a dict, not a set

python
x = {}                          # dict
y = set()                       # set

2. Trying to put a mutable thing in a set

python
# {[1, 2], [3, 4]}    # TypeError — lists aren't hashable
{(1, 2), (3, 4)}      # tuples are fine

3. Indexing a set

python
# s = {"a", "b"}
# s[0]                          # TypeError — sets aren't ordered

If you need indexing, convert to a list first: list(s)[0] (but the "first" item is unpredictable).


Mini-Program — Vocab Overlap Between Two Texts

python
text_a = "the quick brown fox jumps over the lazy dog"
text_b = "the lazy fox sleeps under a brown tree"

words_a = set(text_a.split())
words_b = set(text_b.split())

shared = words_a & words_b
only_a = words_a - words_b
only_b = words_b - words_a

print(f"Shared words: {sorted(shared)}")
print(f"Unique to A:  {sorted(only_a)}")
print(f"Unique to B:  {sorted(only_b)}")
print(f"Vocab overlap: {len(shared) / len(words_a | words_b):.1%}")

That's set math doing real text analysis in 8 lines.


🎯 Your Turn — Who Dropped the Course?

You have two attendance registers: the students enrolled in week 1, and those
still enrolled in week 5. Answer three questions with set operations rather than
loops.

Write course_changes(week1, week5) returning a tuple of three sorted lists:
students who dropped, students who joined late, and students who stayed
throughout.

python
week1 = ["asha", "biju", "chitra", "deepak"]
week5 = ["asha", "chitra", "elango"]

dropped →  ["biju", "deepak"]
joined  →  ["elango"]
stayed  →  ["asha", "chitra"]

Skeleton:

python
def course_changes(week1, week5):
    a = set(week1)
    b = set(week5)
    # TODO 1: dropped = in week1 but not week5
    # TODO 2: joined  = in week5 but not week1
    # TODO 3: stayed  = in both
    # TODO 4: return all three as sorted lists
    ...

dropped, joined, stayed = course_changes(
    ["asha", "biju", "chitra", "deepak"],
    ["asha", "chitra", "elango"],
)
print("dropped:", dropped)   # ['biju', 'deepak']
print("joined :", joined)    # ['elango']
print("stayed :", stayed)    # ['asha', 'chitra']
Hint 1 — Three operators, three answers a - b is everything in a that is not in b. b - a is the reverse. a & b is what appears in both. Each question in this exercise is exactly one of those.
Hint 2 — Sets have no order, so sort on the way out Printing a set gives you an arbitrary order that can change between runs, which makes output impossible to test. sorted(some_set) returns a list in a predictable order — do it once, at the boundary where the data leaves your function.
Show full solution
python
def course_changes(week1, week5):
    a = set(week1)
    b = set(week5)
    return sorted(a - b), sorted(b - a), sorted(a & b)


dropped, joined, stayed = course_changes(
    ["asha", "biju", "chitra", "deepak"],
    ["asha", "chitra", "elango"],
)
print("dropped:", dropped)   # ['biju', 'deepak']
print("joined :", joined)    # ['elango']
print("stayed :", stayed)    # ['asha', 'chitra']

Written with loops this would be three nested passes and about fifteen lines. As
set algebra it is one line, and it reads almost exactly like the question you
were asked — which is the real argument for sets. Note that the answer also
silently handles duplicates in the input registers, because building a set drops
them.


Recap

  • A set is an unordered collection of unique items.
  • Create with {1, 2, 3} or set(...). Use set() for an empty set — {} is a dict.
  • Sets are the fastest way to check membership (x in s) — O(1) lookups.
  • Set math operators: | union, & intersection, - difference, ^ symmetric difference.
  • Use frozenset when you need an immutable, hashable set.

Sets are one of those tools that, once you know them, you'll reach for constantly. "Are these emails unique?" "What overlaps between these two groups?" "Have I seen this user before?" — set, set, set.


Source: adapted from Python official documentation Section 5.4 (Sets). PSF License.

Practice this

on practicepython.in

Short exercises that run in your browser and tell you what your code actually did, not just whether a test passed.