CSV, JSON, and JSONL — Structured Text Done Right
1 · The lesson
readThree formats cover the overwhelming majority of structured text you'll touch in Python: CSV for tabular data, JSON for nested objects, JSONL for streaming. The standard library handles all three competently — and has just enough sharp edges that almost everyone has stubbed a toe on newline="" or json.dumps(datetime.now()) at some point.
This lesson is the working professional's reference for those three formats, plus a quick word on TOML and YAML for when config — not data — is the job.
1. CSV — Reader, Writer, and the One Argument Everyone Forgets
The csv module reads and writes comma-separated (and tab-separated, and pipe-separated) text files. The headline API:
import csv # Read as lists with open("people.csv", encoding="utf-8", newline="") as f: reader = csv.reader(f) header = next(reader) # skip the header row for row in reader: print(row) # ['Alice', '30', 'alice@example.com'] # Read as dicts — almost always what you want with open("people.csv", encoding="utf-8", newline="") as f: for row in csv.DictReader(f): print(row["name"], row["email"])
That newline="" is the famous foot-gun. The csv module does its own newline handling (necessary because fields can contain \n inside quotes), and Python's default text-mode newline translation interferes with it. On Windows especially, omitting newline="" gives you blank rows between every real row. On Linux/Mac it usually works by accident, then your script breaks the day a Windows user runs it.
Rule: every open() of a CSV file gets newline="". Every time.
2. CSV — Writing
The write side mirrors the read side. csv.writer for lists, csv.DictWriter for dicts.
import csv from pathlib import Path rows = [ {"name": "Alice", "age": 30, "email": "alice@example.com"}, {"name": "Bob", "age": 25, "email": "bob@example.com"}, ] with open("out.csv", "w", encoding="utf-8", newline="") as f: writer = csv.DictWriter(f, fieldnames=["name", "age", "email"]) writer.writeheader() writer.writerows(rows) # batches; or .writerow(one)
fieldnames= is explicit by design — you have to declare your schema. That makes the output stable: same columns, same order, every run. writer.writerows(iterable) is faster than calling writerow in a loop and shorter to read.
For data that may contain quotes or commas inside fields, the module handles escaping automatically. Don't write CSVs by hand with f.write(",".join(row)) — you'll get the simple case right and the day a name contains "O'Brien, Jr." your file becomes unparseable.
3. CSV — Quoting, Dialects, and the BOM Gotcha
The csv module's defaults are the "excel" dialect — comma-separated, double-quote for quoting, quote only when necessary. Override per writer:
writer = csv.writer(f, quoting=csv.QUOTE_ALL) # quote every field writer = csv.writer(f, quoting=csv.QUOTE_NONNUMERIC) # quote everything except numbers writer = csv.writer(f, delimiter="\t") # TSV writer = csv.writer(f, delimiter="|") # pipe-separated
setup added so this can run · defines f, csv
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) f = _AutoMock('f') csv = _AutoMock('csv')
For reading files of unknown dialect, the sniffer guesses:
with open("mystery.csv", encoding="utf-8", newline="") as f: sample = f.read(2048) f.seek(0) dialect = csv.Sniffer().sniff(sample) reader = csv.reader(f, dialect)
setup added so this can run · defines csv
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) csv = _AutoMock('csv')
The BOM gotcha — CSVs exported from Excel often start with a Unicode Byte-Order Mark (). With encoding="utf-8" that BOM ends up as the first character of your first field name, and row["Name"] raises KeyError because the key is actually "Name". Fix it with encoding="utf-8-sig":
# Handles BOM (Excel-exported CSVs) transparently with open("excel_export.csv", encoding="utf-8-sig", newline="") as f: for row in csv.DictReader(f): print(row["Name"]) # now works
setup added so this can run · defines csv
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) csv = _AutoMock('csv')
utf-8-sig strips the BOM on read; for write, prefer plain utf-8 (most tooling handles BOM-less UTF-8 fine, and writing a BOM annoys downstream tools).
4. JSON — Four Functions and You're Done
The json module has effectively four functions you need to remember:
| Function | Direction | Operates on |
|---|---|---|
json.dumps(obj) | Python → JSON | string |
json.loads(s) | JSON → Python | string |
json.dump(obj, f) | Python → JSON | file |
json.load(f) | JSON → Python | file |
The s suffix means "string". dump/load work on file objects directly — no need to read into a string first.
import json from pathlib import Path config = {"version": 3, "features": ["dark_mode", "autosave"], "limits": {"max": 100}} # String round-trip text = json.dumps(config) # '{"version": 3, ...}' back = json.loads(text) # {'version': 3, ...} # File round-trip with open("config.json", "w", encoding="utf-8") as f: json.dump(config, f, indent=2) with open("config.json", encoding="utf-8") as f: loaded = json.load(f)
For human-readable output, pass indent=2 (or 4 if you're a sadist) and sort_keys=True for stable diffs in version control:
json.dumps(config, indent=2, sort_keys=True)
setup added so this can run · defines config, json
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) config = _AutoMock('config') json = _AutoMock('json')
Without sort_keys, Python preserves insertion order — fine in code, terrible in a file you intend to diff because key reordering shows up as noise.
5. JSON — Custom Types (datetime, Decimal, sets…)
json.dumps only knows about primitives: dict, list, str, int, float, bool, None. Hand it a datetime, Decimal, set, or any custom class and it raises TypeError: Object of type datetime is not JSON serializable.
Three escape hatches, in increasing order of seriousness:
(a) Convert before encoding — the simplest, fine for one-offs:
from datetime import datetime record = {"name": "Alice", "joined": datetime.now().isoformat()} json.dumps(record)
setup added so this can run · defines json
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) json = _AutoMock('json')
(b) default= callback — one-liner for a single ad-hoc encoder:
from datetime import datetime record = {"name": "Alice", "joined": datetime.now()} json.dumps(record, default=str) # falls back to str() for unknowns
setup added so this can run · defines json
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) json = _AutoMock('json')
default=str works because str(datetime) returns an ISO-like string. For better control, supply your own function:
def encode(o): if isinstance(o, datetime): return o.isoformat() if isinstance(o, set): return sorted(o) raise TypeError(f"can't encode {type(o).__name__}") json.dumps(payload, default=encode)
setup added so this can run · defines payload, datetime, json
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) payload = _AutoMock('payload') datetime = _AutoMock('datetime') json = _AutoMock('json')
(c) Subclass json.JSONEncoder — when you want the same encoding logic everywhere in a project:
class RichEncoder(json.JSONEncoder): def default(self, o): if isinstance(o, datetime): return o.isoformat() if isinstance(o, set): return sorted(o) return super().default(o) json.dumps(payload, cls=RichEncoder, indent=2)
setup added so this can run · defines json, payload, datetime
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) json = _AutoMock('json') payload = _AutoMock('payload') datetime = _AutoMock('datetime')
The reverse direction — turning ISO strings back into datetime — uses object_hook:
from datetime import datetime import re ISO_RE = re.compile(r"^\d{4}-\d{2}-\d{2}T") def parse_dates(obj): for k, v in obj.items(): if isinstance(v, str) and ISO_RE.match(v): try: obj[k] = datetime.fromisoformat(v) except ValueError: pass return obj json.loads('{"joined": "2026-05-14T10:00:00"}', object_hook=parse_dates) # {'joined': datetime.datetime(2026, 5, 14, 10, 0)}
setup added so this can run · defines json
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) json = _AutoMock('json')
object_hook is called for every JSON object as it's decoded — useful for typed parsing, but it runs on every dict so don't make it expensive.
6. JSON Lines (JSONL / NDJSON) — Streaming-Friendly JSON
Plain JSON for collections means one giant array — and you can't read half of it. JSON Lines fixes that: one JSON object per line, no wrapping array, append-friendly.
{"id": 1, "name": "Alice"}
{"id": 2, "name": "Bob"}
{"id": 3, "name": "Carol"}Writing JSONL — one json.dumps per record, then "\n":
import json from pathlib import Path def write_jsonl(path, records): with open(path, "w", encoding="utf-8") as f: for record in records: f.write(json.dumps(record, ensure_ascii=False)) f.write("\n")
Reading JSONL — one json.loads per line, as a generator so memory stays flat:
def read_jsonl(path): with open(path, encoding="utf-8") as f: for line in f: line = line.strip() if line: # skip blank lines defensively yield json.loads(line) for record in read_jsonl("events.jsonl"): process(record)
setup added so this can run · defines process, json
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def process(*_a, **_kw): print('-> process() called') return _AutoMock('process()') json = _AutoMock('json')
ensure_ascii=False is worth knowing — by default json.dumps escapes non-ASCII characters as \u sequences, making "café" become "café". Technically valid, much less readable. Pass ensure_ascii=False and write UTF-8 throughout.
7. When to Use Which
| Use case | Format | Why |
|---|---|---|
| Tabular data, opens in spreadsheets | CSV | Universal, Excel-friendly, flat rows. |
| Nested config / single response | JSON | Native to most APIs, supports nesting. |
| Streaming logs, append-only event data, large datasets | JSONL | Line-by-line reads, easy to tail -f, no array wrapper. |
| Human-edited config files | TOML | Comments, sections, sane string handling. (Section 9) |
| Existing YAML ecosystem (Kubernetes, GitHub Actions) | YAML | Match the surrounding tools. (Section 9) |
A guiding principle: JSON has no comments. People reach for JSON as a config format and then can't annotate it. Use TOML or YAML for human-edited config; reserve JSON for machine-produced/consumed data.
For data that's both tabular and large (multi-GB CSV or JSONL), streaming is non-negotiable — see Section 8.
8. Streaming Big Files — Don't Slurp
The biggest scaling mistake with both CSV and JSON is loading everything into memory. For huge files, stream row-by-row (CSV) or line-by-line (JSONL):
# CSV — DictReader is already a streaming iterator. Don't materialise it. with open("huge.csv", encoding="utf-8", newline="") as f: for row in csv.DictReader(f): process(row) # constant memory # BAD — pulls all 10 million rows into RAM with open("huge.csv", encoding="utf-8", newline="") as f: rows = list(csv.DictReader(f)) # don't
setup added so this can run · defines csv, process
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) csv = _AutoMock('csv') def process(*_a, **_kw): print('-> process() called') return _AutoMock('process()')
For huge JSON (a single multi-GB array), the stdlib json module won't help — it requires the whole array in memory. Reach for ijson (pip install ijson), which streams JSON one item at a time:
import ijson with open("huge.json", "rb") as f: for record in ijson.items(f, "item"): # iterate top-level array entries process(record)
setup added so this can run · defines process
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def process(*_a, **_kw): print('-> process() called') return _AutoMock('process()')
Better yet, store large datasets as JSONL from the start. JSONL streams naturally with the stdlib — one line, one json.loads, done. See generators for the iteration patterns that keep memory flat.
9. TOML and YAML — When Config Beats Data
Python 3.11 added tomllib to the stdlib for reading TOML (the format pyproject.toml uses). It's read-only — no tomllib.dumps. For writing, install tomli-w or tomlkit.
import tomllib from pathlib import Path with open("pyproject.toml", "rb") as f: # NB: binary mode for tomllib config = tomllib.load(f) print(config["project"]["name"])
TOML is the modern default for Python project config — readable, typed, supports comments. Prefer it for new config files unless the surrounding ecosystem (e.g. Kubernetes manifests) is already YAML.
YAML isn't in the stdlib — pip install pyyaml:
import yaml with open("compose.yaml", encoding="utf-8") as f: config = yaml.safe_load(f) # ALWAYS safe_load, never plain load
Use yaml.safe_load, never yaml.load. The latter can deserialise arbitrary Python objects — including code execution — from untrusted YAML. It's the YAML equivalent of pickle.loads on untrusted data. CVEs have been filed; people have been pwned. safe_load is the only correct default.
10. The pandas Cameo
For data-science workflows on tabular data, pandas.read_csv and pandas.read_json are the right tools — they handle dialects, types, missing values, and gigabytes of data with one line:
import pandas as pd df = pd.read_csv("huge.csv", chunksize=10_000) # streams in chunks for chunk in df: process(chunk)
setup added so this can run · defines process
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def process(*_a, **_kw): print('-> process() called') return _AutoMock('process()')
For one-off CSV manipulation or anything that fits in RAM, pandas is faster to write, more expressive, and worth the import. For long-running services where the dependency footprint matters, or for genuine line-at-a-time streaming, the stdlib csv module is the right call.
This is a Python how-to — not a data-science course. A future track covers pandas/Polars properly.
11. Common Mistakes
1. Forgetting newline="" on CSV files
Blank rows on Windows. Always pass newline="" to both open() for read and write.
2. Defaulting the encoding
On Windows, open()'s default is cp1252, not UTF-8 — non-ASCII characters mojibake or crash. Always pass encoding="utf-8" (or "utf-8-sig" for Excel-exported files). See fileio.
3. json.dumps(datetime.now()) → TypeErrorjson doesn't know about datetime, Decimal, set, or your custom classes. Use default=str for quick wins, custom default= callable or JSONEncoder subclass for production.
4. Using JSON as a config format
JSON has no comments. Humans editing config want to annotate. Use TOML (or YAML) for human-edited config; reserve JSON for machine-produced data.
5. list(reader) on huge CSVs
Pulls everything into memory. csv.DictReader is already an iterator — use for row in reader: and stream. Same for JSONL: don't accumulate the records, process them inline.
6. Loading multi-GB JSON arrays with json.load
The stdlib can't stream JSON. If you must read huge JSON, use ijson. Better: produce JSONL instead, where streaming is one line at a time.
7. Hand-writing CSV with ",".join(row)
Fields containing commas, quotes, or newlines will break your file. Use csv.writer — it escapes correctly.
8. yaml.load on untrusted YAML
Arbitrary code execution risk. Always yaml.safe_load. No exceptions.
🎯 Your Turn — Stream a CSV into JSONL
Write convert_csv_to_jsonl(csv_path, jsonl_path) that reads a CSV row by row and writes one JSON object per line to a JSONL file. Constraints:
- O(1) memory — must handle multi-GB inputs without loading them. Streaming only.
- UTF-8 throughout. Handle BOM-prefixed input transparently.
- Use the header row as field names (
csv.DictReaderdoes this for you). ensure_ascii=Falsein the output so non-ASCII characters stay human-readable.
Example:
# input.csv name,age,city Alice,30,Bengaluru Bob,25,München
setup added so this can run · defines name, age, city, Alice, Bengaluru, Bob, München
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) name = _AutoMock('name') age = _AutoMock('age') city = _AutoMock('city') Alice = _AutoMock('Alice') Bengaluru = _AutoMock('Bengaluru') Bob = _AutoMock('Bob') München = _AutoMock('München')
# output.jsonl {"name": "Alice", "age": "30", "city": "Bengaluru"} {"name": "Bob", "age": "25", "city": "München"}
Skeleton:
import csv import json from pathlib import Path def convert_csv_to_jsonl(csv_path, jsonl_path): # TODO 1: open csv with utf-8-sig + newline="" # TODO 2: open jsonl with utf-8 for writing # TODO 3: DictReader on the input # TODO 4: for each row, json.dumps + "\n" to the output ...
Hint 1 — Why utf-8-sig for input
Excel-exported CSVs often start with a UTF-8 BOM (). With plain utf-8, that BOM becomes the first character of your first column header — row["name"] turns into KeyError because the real key is "name". utf-8-sig strips it transparently.
Hint 2 — Streaming write
Don't collect rows into a list first.csv.DictReader is an iterator; iterate it directly, write each row immediately. Memory stays at one row, no matter how big the file is.
Show full solution
import csv import json from pathlib import Path def convert_csv_to_jsonl(csv_path, jsonl_path): """Stream a CSV file row-by-row into a JSONL file. O(1) memory.""" csv_path = Path(csv_path) jsonl_path = Path(jsonl_path) with csv_path.open("r", encoding="utf-8-sig", newline="") as src, \ jsonl_path.open("w", encoding="utf-8") as dst: reader = csv.DictReader(src) for row in reader: dst.write(json.dumps(row, ensure_ascii=False)) dst.write("\n") # Demo import tempfile from pathlib import Path with tempfile.TemporaryDirectory() as tmpdir: tmp = Path(tmpdir) src = tmp / "input.csv" dst = tmp / "output.jsonl" src.write_text( "name,age,city\n" "Alice,30,Bengaluru\n" "Bob,25,München\n", encoding="utf-8", ) convert_csv_to_jsonl(src, dst) print(dst.read_text(encoding="utf-8")) # {"name": "Alice", "age": "30", "city": "Bengaluru"} # {"name": "Bob", "age": "25", "city": "München"}
What's working here:
utf-8-sigon input strips the BOM if present, leaving keys clean ("name", not"name").newline=""lets thecsvmodule handle newlines itself — no blank rows on Windows.csv.DictReaderuses the header row as keys automatically. No hand-built schema needed.- Iteration not materialisation —
for row in reader:reads one row at a time. Memory stays flat for a 10-MB file or a 10-GB file. ensure_ascii=FalsekeepsMünchenasMünchenin the output, notMünchen.- Two files in one
with— both close cleanly even if writing raises mid-stream.
Notes you'd want for production use:
- All values are strings in the output, because CSV has no type information. If you need real types (
"age": 30instead of"age": "30"), add a type-coercion step — usually a per-column mapping like{"age": int, "joined": datetime.fromisoformat}. - No header? No DictReader. Use
csv.readerand build dicts yourself with a known field list. - Errors aren't caught — a malformed row will crash mid-stream. Wrap the inner write in
try/except csv.Errorif you want a forgiving converter that logs and skips bad rows. See exceptions.
The structure — open both, iterate, transform, write line-at-a-time — is the canonical streaming-conversion pattern. The same shape works for JSONL→CSV, CSV→Parquet, log→event-stream. Memorise it once and reuse forever.
What You Learned
- CSV:
csv.DictReader/DictWriter, always withnewline=""andencoding="utf-8"(orutf-8-sigfor BOM-prefixed inputs). - CSV dialects — comma/tab/pipe,
QUOTE_ALL/QUOTE_NONNUMERIC,csv.Snifferfor unknown formats. - JSON: four functions —
dumps/loads/dump/load.indent=2for humans,sort_keys=Truefor diff-friendly output,ensure_ascii=Falsefor UTF-8 readability. - Custom types in JSON:
default=strfor quick fixes,default=callablefor control,JSONEncodersubclass for project-wide encoders. object_hookparses typed values back out on decode (ISO strings →datetime).- JSONL: one JSON object per line. The right format for streaming logs, append-only event data, and anything larger than RAM.
- Format choice: CSV for tabular spreadsheet-friendly, JSON for nested API payloads, JSONL for streams, TOML/YAML for human config.
- Stream, don't slurp.
for row in DictReader(f):andfor line in f:stay O(1). Avoidlist(reader)andjson.loadon multi-GB files. - TOML is stdlib in 3.11+ (read-only). YAML needs
pyyaml— alwayssafe_load, neverload. - pandas for data-science workflows; stdlib for streaming and lean production code.
Next: Automation Scripts — gluing these formats together into the small utilities you actually run daily.
Practice this
on practicepython.inShort exercises that run in your browser and tell you what your code actually did, not just whether a test passed.