UnicodeDecodeError: 'utf-8' codec can't decode byte
1 · The lesson
readWhat this error means
You opened a file (or decoded abytes object) telling Python it was UTF-8, but the bytes don't form valid UTF-8. The codec hit a byte sequence that has no UTF-8 interpretation and refused to guess. UnicodeDecodeError is a subclass of ValueError.
When you see it
Traceback (most recent call last):
File "load.py", line 2, in <module>
text = open("export.csv").read()
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xa3 in position 142:
invalid start byteThe byte and position are the clue. 0xa3 is the pound sign in Windows-1252 / Latin-1; 0xff 0xfe at position 0 is a UTF-16 BOM; 0xef 0xbb 0xbf is a UTF-8 BOM that some tools dislike.
Why it happens
Python 3'sopen() defaults to your platform's text encoding — UTF-8 on modern Linux and macOS, often UTF-8 on Windows 10+ but historically cp1252. A file produced on Windows by Excel ("Save as CSV") is usually cp1252 or has a UTF-8 BOM. A file scraped from an old website may be latin-1. The bytes are valid in their original encoding; they just aren't valid UTF-8.
How to fix it
Option 1 — pass the right encoding. This is the only real fix.
# Excel CSV, Western Europe / US English locale text = open("export.csv", encoding="cp1252").read() # Older web pages, plain text from legacy systems text = open("scrape.html", encoding="latin-1").read() # UTF-8 with a leading BOM () — Excel does this text = open("export.csv", encoding="utf-8-sig").read()
latin-1 is also useful as a fallback because every single byte is a valid latin-1 character — it will never raise — but you may get gibberish for non-Western text.
Option 2 — detect the encoding. Useful when you have many files of unknown provenance.
import chardet # third-party: python -m pip install chardet raw = open("export.csv", "rb").read() guess = chardet.detect(raw) print(guess) # {'encoding': 'Windows-1252', 'confidence': 0.73} text = raw.decode(guess["encoding"])
For UTF-8 vs UTF-16 only, the stdlib codecs module's BOM constants are enough.
Option 3 — tolerate bad bytes. Last resort, when you accept data loss.
text = open("messy.txt", encoding="utf-8", errors="replace").read() # bad bytes become U+FFFD (the replacement character) text = open("messy.txt", encoding="utf-8", errors="ignore").read() # bad bytes silently dropped
errors="surrogateescape" is the right choice if you intend to write the bytes back out untouched — it round-trips invalid bytes through fake code points.
Option 4 — re-export the source as UTF-8. When you control the producer (Excel, a database export), set the output encoding to UTF-8 once and stop fighting this forever.
When you'd actually see this in real code
- "It worked on my Mac" — colleague's CSV exported from Excel on Windows opens fine for them (cp1252 default) and dies on your Linux server (UTF-8 default).
- Scraping HTML that declares
charset=iso-8859-1while you call.textassuming UTF-8. - Reading log files from a 20-year-old appliance that emits Latin-1.
Related errors
UnicodeEncodeError— the reverse direction: you have a string and tried to write it in an encoding that can't represent some character.LookupError: unknown encoding: ...— typo in the encoding name.
See Also
- All Python errors — the full index, by type and by when it happens.
- File I/O — encoding-aware reading and writing.
- CSV and JSON — encoding gotchas with tabular data.
- cheat-encoding — fast lookup.