PythonMastery
intermediate 20 min read · lesson 4 of 12 in Python How-To

CSV, JSON, and JSONL — Structured Text Done Right

1 · The lesson

read

Three formats cover the overwhelming majority of structured text you'll touch in Python: CSV for tabular data, JSON for nested objects, JSONL for streaming. The standard library handles all three competently — and has just enough sharp edges that almost everyone has stubbed a toe on newline="" or json.dumps(datetime.now()) at some point.

This lesson is the working professional's reference for those three formats, plus a quick word on TOML and YAML for when config — not data — is the job.


1. CSV — Reader, Writer, and the One Argument Everyone Forgets

The csv module reads and writes comma-separated (and tab-separated, and pipe-separated) text files. The headline API:

python
import csv

# Read as lists
with open("people.csv", encoding="utf-8", newline="") as f:
    reader = csv.reader(f)
    header = next(reader)                              # skip the header row
    for row in reader:
        print(row)                                     # ['Alice', '30', 'alice@example.com']

# Read as dicts — almost always what you want
with open("people.csv", encoding="utf-8", newline="") as f:
    for row in csv.DictReader(f):
        print(row["name"], row["email"])

That newline="" is the famous foot-gun. The csv module does its own newline handling (necessary because fields can contain \n inside quotes), and Python's default text-mode newline translation interferes with it. On Windows especially, omitting newline="" gives you blank rows between every real row. On Linux/Mac it usually works by accident, then your script breaks the day a Windows user runs it.

Rule: every open() of a CSV file gets newline="". Every time.


2. CSV — Writing

The write side mirrors the read side. csv.writer for lists, csv.DictWriter for dicts.

python
import csv
from pathlib import Path

rows = [
    {"name": "Alice", "age": 30, "email": "alice@example.com"},
    {"name": "Bob",   "age": 25, "email": "bob@example.com"},
]

with open("out.csv", "w", encoding="utf-8", newline="") as f:
    writer = csv.DictWriter(f, fieldnames=["name", "age", "email"])
    writer.writeheader()
    writer.writerows(rows)                             # batches; or .writerow(one)

fieldnames= is explicit by design — you have to declare your schema. That makes the output stable: same columns, same order, every run. writer.writerows(iterable) is faster than calling writerow in a loop and shorter to read.

For data that may contain quotes or commas inside fields, the module handles escaping automatically. Don't write CSVs by hand with f.write(",".join(row)) — you'll get the simple case right and the day a name contains "O'Brien, Jr." your file becomes unparseable.


3. CSV — Quoting, Dialects, and the BOM Gotcha

The csv module's defaults are the "excel" dialect — comma-separated, double-quote for quoting, quote only when necessary. Override per writer:

python
writer = csv.writer(f, quoting=csv.QUOTE_ALL)          # quote every field
writer = csv.writer(f, quoting=csv.QUOTE_NONNUMERIC)   # quote everything except numbers
writer = csv.writer(f, delimiter="\t")                 # TSV
writer = csv.writer(f, delimiter="|")                  # pipe-separated
+ setup added so this can run · defines f, csv
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

f = _AutoMock('f')
csv = _AutoMock('csv')

For reading files of unknown dialect, the sniffer guesses:

python
with open("mystery.csv", encoding="utf-8", newline="") as f:
    sample = f.read(2048)
    f.seek(0)
    dialect = csv.Sniffer().sniff(sample)
    reader = csv.reader(f, dialect)
+ setup added so this can run · defines csv
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

csv = _AutoMock('csv')

The BOM gotcha — CSVs exported from Excel often start with a Unicode Byte-Order Mark (). With encoding="utf-8" that BOM ends up as the first character of your first field name, and row["Name"] raises KeyError because the key is actually "Name". Fix it with encoding="utf-8-sig":

python
# Handles BOM (Excel-exported CSVs) transparently
with open("excel_export.csv", encoding="utf-8-sig", newline="") as f:
    for row in csv.DictReader(f):
        print(row["Name"])                              # now works
+ setup added so this can run · defines csv
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

csv = _AutoMock('csv')

utf-8-sig strips the BOM on read; for write, prefer plain utf-8 (most tooling handles BOM-less UTF-8 fine, and writing a BOM annoys downstream tools).


4. JSON — Four Functions and You're Done

The json module has effectively four functions you need to remember:

FunctionDirectionOperates on
json.dumps(obj)Python → JSONstring
json.loads(s)JSON → Pythonstring
json.dump(obj, f)Python → JSONfile
json.load(f)JSON → Pythonfile

The s suffix means "string". dump/load work on file objects directly — no need to read into a string first.

python
import json
from pathlib import Path

config = {"version": 3, "features": ["dark_mode", "autosave"], "limits": {"max": 100}}

# String round-trip
text = json.dumps(config)                              # '{"version": 3, ...}'
back = json.loads(text)                                # {'version': 3, ...}

# File round-trip
with open("config.json", "w", encoding="utf-8") as f:
    json.dump(config, f, indent=2)

with open("config.json", encoding="utf-8") as f:
    loaded = json.load(f)

For human-readable output, pass indent=2 (or 4 if you're a sadist) and sort_keys=True for stable diffs in version control:

python
json.dumps(config, indent=2, sort_keys=True)
+ setup added so this can run · defines config, json
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

config = _AutoMock('config')
json = _AutoMock('json')

Without sort_keys, Python preserves insertion order — fine in code, terrible in a file you intend to diff because key reordering shows up as noise.


5. JSON — Custom Types (datetime, Decimal, sets…)

json.dumps only knows about primitives: dict, list, str, int, float, bool, None. Hand it a datetime, Decimal, set, or any custom class and it raises TypeError: Object of type datetime is not JSON serializable.

Three escape hatches, in increasing order of seriousness:

(a) Convert before encoding — the simplest, fine for one-offs:

python
from datetime import datetime
record = {"name": "Alice", "joined": datetime.now().isoformat()}
json.dumps(record)
+ setup added so this can run · defines json
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

json = _AutoMock('json')

(b) default= callback — one-liner for a single ad-hoc encoder:

python
from datetime import datetime
record = {"name": "Alice", "joined": datetime.now()}
json.dumps(record, default=str)                        # falls back to str() for unknowns
+ setup added so this can run · defines json
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

json = _AutoMock('json')

default=str works because str(datetime) returns an ISO-like string. For better control, supply your own function:

python
def encode(o):
    if isinstance(o, datetime):
        return o.isoformat()
    if isinstance(o, set):
        return sorted(o)
    raise TypeError(f"can't encode {type(o).__name__}")

json.dumps(payload, default=encode)
+ setup added so this can run · defines payload, datetime, json
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

payload = _AutoMock('payload')
datetime = _AutoMock('datetime')
json = _AutoMock('json')

(c) Subclass json.JSONEncoder — when you want the same encoding logic everywhere in a project:

python
class RichEncoder(json.JSONEncoder):
    def default(self, o):
        if isinstance(o, datetime):
            return o.isoformat()
        if isinstance(o, set):
            return sorted(o)
        return super().default(o)

json.dumps(payload, cls=RichEncoder, indent=2)
+ setup added so this can run · defines json, payload, datetime
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

json = _AutoMock('json')
payload = _AutoMock('payload')
datetime = _AutoMock('datetime')

The reverse direction — turning ISO strings back into datetime — uses object_hook:

python
from datetime import datetime
import re

ISO_RE = re.compile(r"^\d{4}-\d{2}-\d{2}T")

def parse_dates(obj):
    for k, v in obj.items():
        if isinstance(v, str) and ISO_RE.match(v):
            try:
                obj[k] = datetime.fromisoformat(v)
            except ValueError:
                pass
    return obj

json.loads('{"joined": "2026-05-14T10:00:00"}', object_hook=parse_dates)
# {'joined': datetime.datetime(2026, 5, 14, 10, 0)}
+ setup added so this can run · defines json
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

json = _AutoMock('json')

object_hook is called for every JSON object as it's decoded — useful for typed parsing, but it runs on every dict so don't make it expensive.


6. JSON Lines (JSONL / NDJSON) — Streaming-Friendly JSON

Plain JSON for collections means one giant array — and you can't read half of it. JSON Lines fixes that: one JSON object per line, no wrapping array, append-friendly.

python
{"id": 1, "name": "Alice"}
{"id": 2, "name": "Bob"}
{"id": 3, "name": "Carol"}

Writing JSONL — one json.dumps per record, then "\n":

python
import json
from pathlib import Path

def write_jsonl(path, records):
    with open(path, "w", encoding="utf-8") as f:
        for record in records:
            f.write(json.dumps(record, ensure_ascii=False))
            f.write("\n")

Reading JSONL — one json.loads per line, as a generator so memory stays flat:

python
def read_jsonl(path):
    with open(path, encoding="utf-8") as f:
        for line in f:
            line = line.strip()
            if line:                                   # skip blank lines defensively
                yield json.loads(line)

for record in read_jsonl("events.jsonl"):
    process(record)
+ setup added so this can run · defines process, json
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def process(*_a, **_kw):
    print('-> process() called')
    return _AutoMock('process()')
json = _AutoMock('json')

ensure_ascii=False is worth knowing — by default json.dumps escapes non-ASCII characters as \u sequences, making "café" become "café". Technically valid, much less readable. Pass ensure_ascii=False and write UTF-8 throughout.


7. When to Use Which

Use caseFormatWhy
Tabular data, opens in spreadsheetsCSVUniversal, Excel-friendly, flat rows.
Nested config / single responseJSONNative to most APIs, supports nesting.
Streaming logs, append-only event data, large datasetsJSONLLine-by-line reads, easy to tail -f, no array wrapper.
Human-edited config filesTOMLComments, sections, sane string handling. (Section 9)
Existing YAML ecosystem (Kubernetes, GitHub Actions)YAMLMatch the surrounding tools. (Section 9)

A guiding principle: JSON has no comments. People reach for JSON as a config format and then can't annotate it. Use TOML or YAML for human-edited config; reserve JSON for machine-produced/consumed data.

For data that's both tabular and large (multi-GB CSV or JSONL), streaming is non-negotiable — see Section 8.


8. Streaming Big Files — Don't Slurp

The biggest scaling mistake with both CSV and JSON is loading everything into memory. For huge files, stream row-by-row (CSV) or line-by-line (JSONL):

python
# CSV — DictReader is already a streaming iterator. Don't materialise it.
with open("huge.csv", encoding="utf-8", newline="") as f:
    for row in csv.DictReader(f):
        process(row)                                   # constant memory

# BAD — pulls all 10 million rows into RAM
with open("huge.csv", encoding="utf-8", newline="") as f:
    rows = list(csv.DictReader(f))                     # don't
+ setup added so this can run · defines csv, process
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

csv = _AutoMock('csv')
def process(*_a, **_kw):
    print('-> process() called')
    return _AutoMock('process()')

For huge JSON (a single multi-GB array), the stdlib json module won't help — it requires the whole array in memory. Reach for ijson (pip install ijson), which streams JSON one item at a time:

python
import ijson

with open("huge.json", "rb") as f:
    for record in ijson.items(f, "item"):              # iterate top-level array entries
        process(record)
+ setup added so this can run · defines process
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def process(*_a, **_kw):
    print('-> process() called')
    return _AutoMock('process()')

Better yet, store large datasets as JSONL from the start. JSONL streams naturally with the stdlib — one line, one json.loads, done. See generators for the iteration patterns that keep memory flat.


9. TOML and YAML — When Config Beats Data

Python 3.11 added tomllib to the stdlib for reading TOML (the format pyproject.toml uses). It's read-only — no tomllib.dumps. For writing, install tomli-w or tomlkit.

python
import tomllib
from pathlib import Path

with open("pyproject.toml", "rb") as f:                # NB: binary mode for tomllib
    config = tomllib.load(f)

print(config["project"]["name"])

TOML is the modern default for Python project config — readable, typed, supports comments. Prefer it for new config files unless the surrounding ecosystem (e.g. Kubernetes manifests) is already YAML.

YAML isn't in the stdlib — pip install pyyaml:

python
import yaml

with open("compose.yaml", encoding="utf-8") as f:
    config = yaml.safe_load(f)                         # ALWAYS safe_load, never plain load

Use yaml.safe_load, never yaml.load. The latter can deserialise arbitrary Python objects — including code execution — from untrusted YAML. It's the YAML equivalent of pickle.loads on untrusted data. CVEs have been filed; people have been pwned. safe_load is the only correct default.


10. The pandas Cameo

For data-science workflows on tabular data, pandas.read_csv and pandas.read_json are the right tools — they handle dialects, types, missing values, and gigabytes of data with one line:

python
import pandas as pd

df = pd.read_csv("huge.csv", chunksize=10_000)         # streams in chunks
for chunk in df:
    process(chunk)
+ setup added so this can run · defines process
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def process(*_a, **_kw):
    print('-> process() called')
    return _AutoMock('process()')

For one-off CSV manipulation or anything that fits in RAM, pandas is faster to write, more expressive, and worth the import. For long-running services where the dependency footprint matters, or for genuine line-at-a-time streaming, the stdlib csv module is the right call.

This is a Python how-to — not a data-science course. A future track covers pandas/Polars properly.


11. Common Mistakes

1. Forgetting newline="" on CSV files
Blank rows on Windows. Always pass newline="" to both open() for read and write.

2. Defaulting the encoding
On Windows, open()'s default is cp1252, not UTF-8 — non-ASCII characters mojibake or crash. Always pass encoding="utf-8" (or "utf-8-sig" for Excel-exported files). See fileio.

3. json.dumps(datetime.now()) → TypeError
json doesn't know about datetime, Decimal, set, or your custom classes. Use default=str for quick wins, custom default= callable or JSONEncoder subclass for production.

4. Using JSON as a config format
JSON has no comments. Humans editing config want to annotate. Use TOML (or YAML) for human-edited config; reserve JSON for machine-produced data.

5. list(reader) on huge CSVs
Pulls everything into memory. csv.DictReader is already an iterator — use for row in reader: and stream. Same for JSONL: don't accumulate the records, process them inline.

6. Loading multi-GB JSON arrays with json.load
The stdlib can't stream JSON. If you must read huge JSON, use ijson. Better: produce JSONL instead, where streaming is one line at a time.

7. Hand-writing CSV with ",".join(row)
Fields containing commas, quotes, or newlines will break your file. Use csv.writer — it escapes correctly.

8. yaml.load on untrusted YAML
Arbitrary code execution risk. Always yaml.safe_load. No exceptions.


🎯 Your Turn — Stream a CSV into JSONL

Write convert_csv_to_jsonl(csv_path, jsonl_path) that reads a CSV row by row and writes one JSON object per line to a JSONL file. Constraints:

  • O(1) memory — must handle multi-GB inputs without loading them. Streaming only.
  • UTF-8 throughout. Handle BOM-prefixed input transparently.
  • Use the header row as field names (csv.DictReader does this for you).
  • ensure_ascii=False in the output so non-ASCII characters stay human-readable.

Example:

python
# input.csv
name,age,city
Alice,30,Bengaluru
Bob,25,München
+ setup added so this can run · defines name, age, city, Alice, Bengaluru, Bob, München
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

name = _AutoMock('name')
age = _AutoMock('age')
city = _AutoMock('city')
Alice = _AutoMock('Alice')
Bengaluru = _AutoMock('Bengaluru')
Bob = _AutoMock('Bob')
München = _AutoMock('München')
python
# output.jsonl
{"name": "Alice", "age": "30", "city": "Bengaluru"}
{"name": "Bob", "age": "25", "city": "München"}

Skeleton:

python
import csv
import json
from pathlib import Path

def convert_csv_to_jsonl(csv_path, jsonl_path):
    # TODO 1: open csv with utf-8-sig + newline=""
    # TODO 2: open jsonl with utf-8 for writing
    # TODO 3: DictReader on the input
    # TODO 4: for each row, json.dumps + "\n" to the output
    ...
Hint 1 — Why utf-8-sig for input Excel-exported CSVs often start with a UTF-8 BOM (). With plain utf-8, that BOM becomes the first character of your first column header — row["name"] turns into KeyError because the real key is "name". utf-8-sig strips it transparently.
Hint 2 — Streaming write Don't collect rows into a list first. csv.DictReader is an iterator; iterate it directly, write each row immediately. Memory stays at one row, no matter how big the file is.
Show full solution
python
import csv
import json
from pathlib import Path

def convert_csv_to_jsonl(csv_path, jsonl_path):
    """Stream a CSV file row-by-row into a JSONL file. O(1) memory."""
    csv_path = Path(csv_path)
    jsonl_path = Path(jsonl_path)

    with csv_path.open("r", encoding="utf-8-sig", newline="") as src, \
         jsonl_path.open("w", encoding="utf-8") as dst:
        reader = csv.DictReader(src)
        for row in reader:
            dst.write(json.dumps(row, ensure_ascii=False))
            dst.write("\n")


# Demo
import tempfile
from pathlib import Path

with tempfile.TemporaryDirectory() as tmpdir:
    tmp = Path(tmpdir)
    src = tmp / "input.csv"
    dst = tmp / "output.jsonl"

    src.write_text(
        "name,age,city\n"
        "Alice,30,Bengaluru\n"
        "Bob,25,München\n",
        encoding="utf-8",
    )

    convert_csv_to_jsonl(src, dst)
    print(dst.read_text(encoding="utf-8"))
    # {"name": "Alice", "age": "30", "city": "Bengaluru"}
    # {"name": "Bob", "age": "25", "city": "München"}

What's working here:

  • utf-8-sig on input strips the BOM if present, leaving keys clean ("name", not "name").
  • newline="" lets the csv module handle newlines itself — no blank rows on Windows.
  • csv.DictReader uses the header row as keys automatically. No hand-built schema needed.
  • Iteration not materialisation — for row in reader: reads one row at a time. Memory stays flat for a 10-MB file or a 10-GB file.
  • ensure_ascii=False keeps München as München in the output, not München.
  • Two files in one with — both close cleanly even if writing raises mid-stream.

Notes you'd want for production use:

  • All values are strings in the output, because CSV has no type information. If you need real types ("age": 30 instead of "age": "30"), add a type-coercion step — usually a per-column mapping like {"age": int, "joined": datetime.fromisoformat}.
  • No header? No DictReader. Use csv.reader and build dicts yourself with a known field list.
  • Errors aren't caught — a malformed row will crash mid-stream. Wrap the inner write in try/except csv.Error if you want a forgiving converter that logs and skips bad rows. See exceptions.

The structure — open both, iterate, transform, write line-at-a-time — is the canonical streaming-conversion pattern. The same shape works for JSONL→CSV, CSV→Parquet, log→event-stream. Memorise it once and reuse forever.


What You Learned

  • CSV: csv.DictReader/DictWriter, always with newline="" and encoding="utf-8" (or utf-8-sig for BOM-prefixed inputs).
  • CSV dialects — comma/tab/pipe, QUOTE_ALL / QUOTE_NONNUMERIC, csv.Sniffer for unknown formats.
  • JSON: four functions — dumps/loads/dump/load. indent=2 for humans, sort_keys=True for diff-friendly output, ensure_ascii=False for UTF-8 readability.
  • Custom types in JSON: default=str for quick fixes, default=callable for control, JSONEncoder subclass for project-wide encoders.
  • object_hook parses typed values back out on decode (ISO strings → datetime).
  • JSONL: one JSON object per line. The right format for streaming logs, append-only event data, and anything larger than RAM.
  • Format choice: CSV for tabular spreadsheet-friendly, JSON for nested API payloads, JSONL for streams, TOML/YAML for human config.
  • Stream, don't slurp. for row in DictReader(f): and for line in f: stay O(1). Avoid list(reader) and json.load on multi-GB files.
  • TOML is stdlib in 3.11+ (read-only). YAML needs pyyaml — always safe_load, never load.
  • pandas for data-science workflows; stdlib for streaming and lean production code.

Next: Automation Scripts — gluing these formats together into the small utilities you actually run daily.

Practice this

on practicepython.in

Short exercises that run in your browser and tell you what your code actually did, not just whether a test passed.