File I/O, Pathlib, and Large Files
1 · The lesson
readYou already know with open(...) as f, the read/write modes, and that for line in f beats f.read() for anything non-trivial. This lesson is what professional Python file code actually looks like — pathlib instead of string-mangled paths, streaming for huge files, atomic writes that survive crashes, and the encoding gotcha that gets every cross-platform shop exactly once.
If os.path.join still feels like the right answer, the next 18 minutes will retire it.
1. Pathlib — Stop Treating Paths as Strings
pathlib.Path is the object-oriented path API. A Path knows its parent, its suffix, its parts; it can read, write, glob, and stat itself. You almost never need raw strings or os.path again.
from pathlib import Path p = Path("data") / "logs" / "app.log" # the / operator builds paths print(p) # data/logs/app.log (or data\logs\app.log on Windows) print(p.name) # app.log print(p.stem) # app print(p.suffix) # .log print(p.parent) # data/logs print(p.parts) # ('data', 'logs', 'app.log')
The / operator is overloaded on Path. It does the right thing on every platform, with no os.sep ceremony.
2. The Five Pathlib Methods You'll Use Daily
from pathlib import Path root = Path("project") root.mkdir(parents=True, exist_ok=True) # mkdir -p; no error if it exists (root / "config.yaml").exists() # True/False (root / "config.yaml").is_file() # True only if it exists AND is a file (root / "src").is_dir() # likewise for directories # Read/write in one call — no `with` ceremony for short files (root / "README.md").write_text("# Hi\n", encoding="utf-8") content = (root / "README.md").read_text(encoding="utf-8") # Binary equivalents (root / "thumb.png").write_bytes(b"\x89PNG...") blob = (root / "thumb.png").read_bytes()
A couple more that earn their keep:
# Globbing — find everything matching a pattern for csv in Path("data").glob("*.csv"): print(csv) # Recursive glob — ** crosses subdirectories for py in Path("src").rglob("*.py"): print(py) # Suffix swap — common when transforming files src = Path("notes.md") out = src.with_suffix(".html") # notes.html # Iterate a directory's direct children for child in Path(".").iterdir(): print(child, "(dir)" if child.is_dir() else "(file)")
setup added so this can run · defines Path
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def Path(*_a, **_kw): print('-> Path() called') return _AutoMock('Path()')
write_text and read_text are fine for small files where you genuinely want the whole thing in memory. For anything that could grow, keep using with open(...) and stream it (see Section 5).
3. Text vs Binary Mode — When Bytes Matter
Text mode ("r", "w", "a") decodes bytes to str on read and encodes on write. Binary mode ("rb", "wb", "ab") gives you raw bytes — no decoding, no newline translation.
Use binary mode for anything that isn't human-readable text:
# Hashing a file — bytes only, never decode import hashlib def file_sha256(path): h = hashlib.sha256() with open(path, "rb") as f: for chunk in iter(lambda: f.read(8192), b""): h.update(chunk) return h.hexdigest()
Common binary cases: images, audio, video, pickle files, executables, PDFs, anything you found out is binary by getting a UnicodeDecodeError. If you try to open a JPEG in text mode, Python will guess an encoding, fail to decode some byte, and raise.
On Windows, text mode also translates \n to \r\n on write. That's invisible most of the time and catastrophic when you're computing a hash or comparing file sizes. When in doubt, binary.
4. Encoding — The "Works on My Machine" Trap
Text files are bytes plus an encoding. The world has standardised on UTF-8. Most operating systems default to UTF-8. Windows, until very recently, defaulted Python's open() to cp1252 — a Western European code page that can't represent most non-ASCII text.
# DANGER — encoding defaults to the OS's preferred encoding with open("notes.txt", "w") as f: f.write("café — 北京 — 🐍") # crashes on Windows cp1252 # CORRECT — always specify, every time with open("notes.txt", "w", encoding="utf-8") as f: f.write("café — 北京 — 🐍")
The failure mode is the worst kind: works on your Mac, works in CI on Linux, then a Windows user runs it and gets UnicodeEncodeError: 'charmap' codec can't encode character. Or — worse — silently writes mojibake.
Rules:
- Reading text? Pass
encoding="utf-8". - Writing text? Pass
encoding="utf-8". - Reading text from a source that genuinely uses something else (legacy logs, a CSV from a Windows export)? Pass that encoding explicitly.
- Never trust the default. It's two extra words.
For paranoid reads where the source might be slightly broken:
with open("legacy.txt", "r", encoding="utf-8", errors="replace") as f: text = f.read() # bad bytes become � instead of crashing
5. Reading Huge Files — Iterate, Don't Slurp
The single most common file-handling bug at scale is loading a 4-GB log into memory. The fix is the iteration pattern you already know — but it's worth seeing the contrast explicitly.
# BAD — loads the entire file into a list of strings with open("server.log", encoding="utf-8") as f: lines = f.readlines() # 4 GB file → 4 GB+ in RAM for line in lines: process(line) # GOOD — streams one line at a time, constant memory with open("server.log", encoding="utf-8") as f: for line in f: # file objects are iterators process(line)
setup added so this can run · defines process
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def process(*_a, **_kw): print('-> process() called') return _AutoMock('process()')
The file object is its own iterator. Each next() call reads up to the next \n and yields the line (including the trailing \n). Memory usage stays at one buffer's worth, no matter how big the file is.
Need only the first line? f.readline(). Need only the last ten? Use collections.deque with a max length — it keeps the most recent items as you stream:
from collections import deque with open("server.log", encoding="utf-8") as f: last_ten = deque(f, maxlen=10) # tail -10, in O(file_size) time and O(10) memory for line in last_ten: print(line, end="")
6. Streaming Binary Chunks
For binary files, you stream by chunk size rather than line. The canonical loop uses the walrus operator (:=) — Python's assignment expression:
with open("video.mp4", "rb") as f: while chunk := f.read(8192): # read 8 KB; stop when read returns b"" process(chunk)
setup added so this can run · defines process, chunk
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def process(*_a, **_kw): print('-> process() called') return _AutoMock('process()') chunk = _AutoMock('chunk')
chunk := f.read(8192) assigns and evaluates. The loop ends when f.read() returns b"" (empty bytes), which is the EOF signal in binary mode. 8 KB is a fine default; 64 KB to 1 MB is better for high-throughput work on SSDs.
The pre-walrus version is uglier but works on Python ≤ 3.7:
with open("video.mp4", "rb") as f: for chunk in iter(lambda: f.read(8192), b""): process(chunk)
setup added so this can run · defines process
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def process(*_a, **_kw): print('-> process() called') return _AutoMock('process()')
iter(callable, sentinel) keeps calling callable() until it returns the sentinel. Both forms are idiomatic.
7. Atomic Writes — Surviving Crashes
The naive pattern has a window where the file on disk is half-written:
# BAD — if Python crashes mid-write, config.json is corrupt with open("config.json", "w", encoding="utf-8") as f: f.write(serialize(big_object)) # crash here → empty/partial file
setup added so this can run · defines serialize, big_object
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def serialize(*_a, **_kw): print('-> serialize() called') return _AutoMock('serialize()') big_object = _AutoMock('big_object')
The professional pattern is write to a temp file, then rename:
import os from pathlib import Path def atomic_write_text(path, text, encoding="utf-8"): """Write text to `path` atomically — readers see either the old file or the full new one.""" path = Path(path) tmp = path.with_suffix(path.suffix + ".tmp") tmp.write_text(text, encoding=encoding) os.replace(tmp, path) # atomic on every major OS
os.replace() is atomic on POSIX and Windows — the rename either fully happens or doesn't. Readers of the target file never see a half-written state. Configuration files, sqlite-style databases, and anything that other processes might read while you're writing should use this pattern.
For extra paranoia (crash between write and replace), call tmp_fh.flush() and os.fsync(tmp_fh.fileno()) before closing the temp file. That forces the OS to push the bytes to disk, not just to its buffer cache. Most apps don't need this; databases do.
8. Temporary Files and Directories
The tempfile module gives you safe, OS-chosen temp paths that get cleaned up automatically.
import tempfile from pathlib import Path # A one-off temp FILE — auto-deleted when the block exits with tempfile.NamedTemporaryFile(mode="w", encoding="utf-8", suffix=".json", delete_on_close=False) as tmp: tmp.write('{"hello": "world"}') tmp.flush() print("Temp path:", tmp.name) # use tmp.name as a real path here — pass it to another tool, etc. # file is unlinked here # A whole temp DIRECTORY — auto-removed (recursively) when the block exits with tempfile.TemporaryDirectory() as tmpdir: workdir = Path(tmpdir) (workdir / "input.txt").write_text("hi", encoding="utf-8") # ... do work in workdir ... # directory and contents are gone
NamedTemporaryFile is the right tool whenever you need a real on-disk path to hand to a subprocess or another library. TemporaryDirectory is the right tool for test fixtures, build sandboxes, and any "I need to write a tree of files and throw it away."
Both rely on with-block cleanup — we'll go deeper on building your own context managers in Advanced.
9. shutil Highlights
shutil is the standard library's high-level file operation toolkit. The four you'll actually use:
import shutil from pathlib import Path shutil.copy("src.txt", "dst.txt") # copy a single file (metadata if .copy2) shutil.copytree("src_dir", "dst_dir") # copy a directory tree shutil.move("from.txt", "to.txt") # rename/move across filesystems shutil.rmtree("scratch_dir") # rm -rf — irreversible, no Trash usage = shutil.disk_usage(Path.home()) print(f"{usage.free / 1e9:.1f} GB free")
shutil.rmtree is a chainsaw. There's no undo. Double-check the path, and consider ignore_errors=False (the default) so you actually see when something failed.
10. CSV and JSON — A One-Liner Intro
You'll often want structured text rather than raw lines. The standard library has both:
import csv, json # CSV — read as rows of dicts with open("people.csv", encoding="utf-8", newline="") as f: for row in csv.DictReader(f): print(row["name"], row["email"]) # JSON — read a whole file into a Python object with open("config.json", encoding="utf-8") as f: config = json.load(f) # JSON — write a Python object back out with open("config.json", "w", encoding="utf-8") as f: json.dump(config, f, indent=2)
The newline="" argument on CSV files is required — csv does its own newline handling, and the default text-mode translation interferes with quoted fields. Both modules have a full lesson's worth of options (dialects, custom encoders, streaming); the CSV & JSON How-To covers them.
11. Common Mistakes
1. Forgetting encoding=
Works on your machine, breaks for everyone else. Adding encoding="utf-8" is the cheapest reliability win in Python. Every text open() should have it.
2. Holding file handles open longer than needed
If you're not using with, you're doing it wrong. Holding a handle past its useful life ties up OS file descriptors, locks the file on Windows, and risks data loss if your process is killed before the buffer flushes.
3. Reading the whole file when you only need partf.read() for a 4 GB file when you only wanted the first line is the rookie mistake. f.readline() for the first line; collections.deque(f, maxlen=N) for the last N; for line in f for everything in between.
4. Comparing paths as strings"data/log.txt" == "data\\log.txt" is False. Path("data/log.txt") == Path("data\\log.txt") is True on Windows. Once you've got Path objects, compare and hash them as paths — don't str() them just to use ==.
5. The check-then-act race
# RACE CONDITION — another process could create the file between these two lines if not path.exists(): path.write_text("default")
setup added so this can run · defines path
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) path = _AutoMock('path')
Between the
exists() and the write_text(), another process or thread could create the file. Either use exclusive-create mode (open(path, "x") raises FileExistsError if it exists), use Path.write_text(...) unconditionally if you're happy to overwrite, or catch the exception:
try: with open(path, "x", encoding="utf-8") as f: f.write("default") except FileExistsError: pass # someone else created it; that's fine
setup added so this can run · defines path
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) path = _AutoMock('path')
This is the EAFP principle in action — easier to ask forgiveness than permission. We'll cover it more in Exceptions.
🎯 Your Turn — Stream-Dedupe a Huge File
Write dedupe_lines(input_path, output_path) that reads a (potentially huge) text file and writes the unique lines, in first-seen order, to the output file. The constraint: memory must be O(unique line count), not O(file size). You can't slurp the file with readlines().
# Input file (one line per row): # alpha # beta # alpha # gamma # beta # delta # Output file: # alpha # beta # gamma # delta
Skeleton:
from pathlib import Path def dedupe_lines(input_path, output_path): # TODO 1: open input for reading, output for writing — both UTF-8 # TODO 2: keep a set of lines you've already seen # TODO 3: stream input line-by-line; if a line is new, write it and add to the set ...
Hint 1 — Streaming, not slurping
Iterate the file object directly:for line in f:. Each iteration reads one line. Don't call readlines() — that loads everything.
Hint 2 — Memory-efficient deduplication
Aset gives you O(1) membership checks. You store each unique line in it the first time you see it. The set grows with the number of distinct lines, not the total lines — which is what the constraint asks for.
Show full solution
from pathlib import Path def dedupe_lines(input_path, output_path): """Stream input, write only first-occurrence lines. O(unique) memory.""" seen = set() in_path = Path(input_path) out_path = Path(output_path) with in_path.open("r", encoding="utf-8") as src, \ out_path.open("w", encoding="utf-8") as dst: for line in src: if line not in seen: seen.add(line) dst.write(line) # Demo import tempfile from pathlib import Path with tempfile.TemporaryDirectory() as tmpdir: tmp = Path(tmpdir) src = tmp / "in.txt" dst = tmp / "out.txt" src.write_text("alpha\nbeta\nalpha\ngamma\nbeta\ndelta\n", encoding="utf-8") dedupe_lines(src, dst) print(dst.read_text(encoding="utf-8")) # alpha # beta # gamma # delta
What you did:
- Streamed the input —
for line in srcreads one line at a time, never loading the whole file. - Used a
setfor membership — O(1) check, O(unique) storage. - Wrote each new line straight through — no intermediate list.
- Wrapped both files in one
withso they close even if writing fails halfway.
If the input had 10 million lines but only 100 unique, your memory usage is 100 entries, not 10 million. That's the difference between a script that runs and a script that gets OOM-killed at 3 a.m.
For true unbounded uniqueness on massive inputs where even the unique set won't fit, you'd need an external-memory dedup (sort + uniq, or a Bloom filter for probabilistic dedup). That's outside this lesson — but the streaming pattern is the foundation.
What You Learned
- Pathlib replaces
os.path. Build paths with/, query with.exists()/.is_file()/.glob(), mutate with.with_suffix(), create with.mkdir(parents=True, exist_ok=True). - Binary mode (
"rb","wb") for anything not human-readable text — hashes, images, pickles. Text mode silently corrupts non-UTF-8 bytes. - Always pass
encoding="utf-8"— defaulting bites you on Windows. - Stream, don't slurp.
for line in ffor text,while chunk := f.read(8192)for binary. O(1) memory. - Atomic writes: write to
*.tmp, thenos.replace(tmp, final). Crash-safe. tempfile.NamedTemporaryFileandTemporaryDirectoryclean up after themselves.shutilfor copy/move/rmtree/disk_usage.rmtreeis permanent — measure twice.csvandjsonfor structured text — full treatment in the CSV & JSON How-To.- Avoid check-then-act races — use
open(path, "x")or catch the exception (EAFP).
Next: Exceptions: Patterns Beyond the Basics — custom exception hierarchies, chaining, ExceptionGroup, and the patterns production code actually relies on.
Practice this
on practicepython.inShort exercises that run in your browser and tell you what your code actually did, not just whether a test passed.