dict.fromkeys keeps the first time each item appears, in order.
emails = ["ada@x.com", "linus@x.com", "ada@x.com", "grace@x.com", "linus@x.com"] unique = list(dict.fromkeys(emails)) print(unique)
['ada@x.com', 'linus@x.com', 'grace@x.com']
It works because dictionary keys are unique and, since Python 3.7, stay in the order they were added. You build a dictionary just for its keys, then take them back out as a list.
Why not set()?
list(set(emails)) also removes duplicates, but a set has no order. You get the items back in whatever order Python finds convenient:
print(sorted(set([3, 1, 3, 2, 1])))
[1, 2, 3]
Use a set when order doesn't matter, or when you're going to sort anyway, as above. Use dict.fromkeys when the order means something, like a history or a queue.
Duplicates that aren't exactly equal
"Ada@X.com" and "ada@x.com" are different strings. To de-duplicate by a rule, track what you've seen in a set and keep the first original:
emails = ["Ada@X.com", "linus@x.com", "ada@x.com ", "GRACE@x.com"] seen, unique = set(), [] for email in emails: key = email.strip().lower() if key not in seen: seen.add(key) unique.append(email) print(unique)
['Ada@X.com', 'linus@x.com', 'GRACE@x.com']
Lists of dictionaries
Both quick methods fail on a list of dicts, with TypeError: unhashable type: 'dict': dictionaries can't be set members or dictionary keys. Use the same seen-set pattern, keyed on the field that defines "the same":
orders = [
{"id": "A-1", "total": 10},
{"id": "A-2", "total": 25},
{"id": "A-1", "total": 10},
]
seen, unique = set(), []
for order in orders:
if order["id"] not in seen:
seen.add(order["id"])
unique.append(order)
print([o["id"] for o in unique])['A-1', 'A-2']
The unhashable type error page explains why lists and dicts can't go in a set.