PythonMastery

Check membership in a set, not a list

x in some_list checks items one by one; x in some_set jumps straight to the answer. Measure the difference yourself, in the browser.

x in items reads the same for a list and a set, but they do very different work. A list is checked item by item. A set hashes x and goes straight to where it would be.

Measure it

Run this. The numbers are yours, from your own machine.

python
import timeit

ids_list = list(range(100_000))
ids_set = set(ids_list)

# Per lookup, and enough repeats that even a browser's coarse clock sees the set.
per_list = timeit.timeit(lambda: 99_999 in ids_list, number=200) / 200
per_set = timeit.timeit(lambda: 99_999 in ids_set, number=200_000) / 200_000

print(f"list: {per_list * 1e6:,.1f} µs per lookup   set: {per_set * 1e6:.3f} µs per lookup")
print(f"the set was about {per_list / per_set:,.0f}x faster")
example output · yours will differ
list: 643.0 µs per lookup   set: 0.046 µs per lookup
the set was about 14,067x faster

The gap grows with the data: a list twice as long takes twice as long to search, while the set barely notices.

The pattern in real code

The usual shape is "keep the rows whose ID is in some other collection". Build the set once, outside the loop:

python
orders = [("A-1", 30), ("A-2", 12), ("A-3", 55), ("A-4", 8)]
flagged = ["A-2", "A-4"]

flagged_ids = set(flagged)          # once
safe = [o for o in orders if o[0] not in flagged_ids]
print(safe)
output
[('A-1', 30), ('A-3', 55)]

When not to use it

Building a set reads every item once, so for a single lookup a list is just as quick. A set also keeps no order and no duplicates; if either matters, keep the list and build a set alongside it for the lookups.

Learn it properly: Sets: Unordered, Unique, Lightning Fast, Performance: Profile, Then Optimise