PythonMastery
reference 3 min read · lesson 28 of 45 in Errors

UnicodeDecodeError: 'utf-8' codec can't decode byte

1 · The lesson

read

What this error means

You opened a file (or decoded a bytes object) telling Python it was UTF-8, but the bytes don't form valid UTF-8. The codec hit a byte sequence that has no UTF-8 interpretation and refused to guess. UnicodeDecodeError is a subclass of ValueError.

When you see it

text
Traceback (most recent call last):
  File "load.py", line 2, in <module>
    text = open("export.csv").read()
UnicodeDecodeError: 'utf-8' codec can't decode byte 0xa3 in position 142:
invalid start byte

The byte and position are the clue. 0xa3 is the pound sign in Windows-1252 / Latin-1; 0xff 0xfe at position 0 is a UTF-16 BOM; 0xef 0xbb 0xbf is a UTF-8 BOM that some tools dislike.

Why it happens

Python 3's open() defaults to your platform's text encoding — UTF-8 on modern Linux and macOS, often UTF-8 on Windows 10+ but historically cp1252. A file produced on Windows by Excel ("Save as CSV") is usually cp1252 or has a UTF-8 BOM. A file scraped from an old website may be latin-1. The bytes are valid in their original encoding; they just aren't valid UTF-8.

How to fix it

Option 1 — pass the right encoding. This is the only real fix.

python
# Excel CSV, Western Europe / US English locale
text = open("export.csv", encoding="cp1252").read()

# Older web pages, plain text from legacy systems
text = open("scrape.html", encoding="latin-1").read()

# UTF-8 with a leading BOM () — Excel does this
text = open("export.csv", encoding="utf-8-sig").read()

latin-1 is also useful as a fallback because every single byte is a valid latin-1 character — it will never raise — but you may get gibberish for non-Western text.

Option 2 — detect the encoding. Useful when you have many files of unknown provenance.

python
import chardet            # third-party: python -m pip install chardet

raw = open("export.csv", "rb").read()
guess = chardet.detect(raw)
print(guess)              # {'encoding': 'Windows-1252', 'confidence': 0.73}
text = raw.decode(guess["encoding"])

For UTF-8 vs UTF-16 only, the stdlib codecs module's BOM constants are enough.

Option 3 — tolerate bad bytes. Last resort, when you accept data loss.

python
text = open("messy.txt", encoding="utf-8", errors="replace").read()
# bad bytes become U+FFFD (the replacement character)

text = open("messy.txt", encoding="utf-8", errors="ignore").read()
# bad bytes silently dropped

errors="surrogateescape" is the right choice if you intend to write the bytes back out untouched — it round-trips invalid bytes through fake code points.

Option 4 — re-export the source as UTF-8. When you control the producer (Excel, a database export), set the output encoding to UTF-8 once and stop fighting this forever.

When you'd actually see this in real code

  • "It worked on my Mac" — colleague's CSV exported from Excel on Windows opens fine for them (cp1252 default) and dies on your Linux server (UTF-8 default).
  • Scraping HTML that declares charset=iso-8859-1 while you call .text assuming UTF-8.
  • Reading log files from a 20-year-old appliance that emits Latin-1.
  • UnicodeEncodeError — the reverse direction: you have a string and tried to write it in an encoding that can't represent some character.
  • LookupError: unknown encoding: ... — typo in the encoding name.

See Also