Python for Data Analysis
1 · The lesson
readBefore NumPy and pandas arrive, Python's standard library already does a surprising amount of analysis work. This lesson covers the DS-flavoured idioms — comprehensions, zip, Counter, generators — and the file formats you'll meet on day one. By the end you'll have written a tiny analytics script with zero third-party imports.
1. Comprehensions Are the Lingua Franca
A data analysis script transforms one collection into another. Comprehensions express that in one line. You met them in Comprehensions; here's how DS code uses them.
prices = [9.99, 14.50, 7.25, 22.00, 3.99] # Transform — add VAT to each price with_vat = [round(p * 1.2, 2) for p in prices] print(with_vat) # [11.99, 17.4, 8.7, 26.4, 4.79] # Filter — only items above £10 expensive = [p for p in prices if p > 10] print(expensive) # [14.5, 22.0] # Map to a dict sku_to_price = {f"SKU-{i:03d}": p for i, p in enumerate(prices, start=1)} print(sku_to_price) # {'SKU-001': 9.99, 'SKU-002': 14.5, 'SKU-003': 7.25, 'SKU-004': 22.0, 'SKU-005': 3.99} # Aggregate — total revenue if we sold one of each total = sum(p for p in prices) print(f"Total: £{total:.2f}") # Total: £57.73
Memorise the shapes: [expr for x in iter if cond], {k: v for ...}, sum(expr for ...). They cover most cleanup tasks before pandas enters the room.
2. enumerate and zip — The Pair You Actually Use
enumerate gives you (index, value) as you loop. zip walks several iterables in lockstep. Together they replace 80% of range(len(...))-style code.
products = ["Coffee", "Tea", "Cocoa"] prices = [3.50, 2.75, 4.00] units = [120, 340, 85] # enumerate — numbered output for i, name in enumerate(products, start=1): print(f"{i}. {name}") # zip — pair columns together revenue = {name: price * qty for name, price, qty in zip(products, prices, units)} print(revenue) # {'Coffee': 420.0, 'Tea': 935.0, 'Cocoa': 340.0}
zip stops at the shortest iterable — pass strict=True (Python 3.10+) if you want it to raise when lengths mismatch.
3. collections.Counter — Tallies in One Line
Whenever you'd write if key in d: d[key] += 1 else: d[key] = 1, reach for Counter.
from collections import Counter orders = ["coffee", "tea", "coffee", "coffee", "cocoa", "tea", "coffee"] tally = Counter(orders) print(tally) # Counter({'coffee': 4, 'tea': 2, 'cocoa': 1}) print(tally.most_common(2)) # [('coffee', 4), ('tea', 2)] print(tally["coffee"]) # 4 print(tally["water"]) # 0 — missing keys return 0, never raise
most_common() is the one-line "top-N categories" function you'll write a thousand times. Counter also supports arithmetic — Counter(a) + Counter(b) merges two tallies.
4. Loading CSV with the Standard Library
Most DS data arrives as CSV. The stdlib csv module handles it without any dependency. (Recap and details in CSV & JSON.)
import csv from io import StringIO # A tiny CSV inline so we don't need a file csv_text = """product,price,quantity Coffee,3.50,120 Tea,2.75,340 Cocoa,4.00,85 """ # DictReader yields each row as {column: value} reader = csv.DictReader(StringIO(csv_text)) rows = list(reader) for row in rows: print(row) # {'product': 'Coffee', 'price': '3.50', 'quantity': '120'} # ...
Note the values are strings — CSV has no types. You convert them yourself:
typed = [
{"product": r["product"], "price": float(r["price"]), "quantity": int(r["quantity"])}
for r in rows
] setup added so this can run · defines rows
rows = [{"product": "alpha", "price": "3", "quantity": "3"}, {"product": "beta", "price": "6", "quantity": "6"}, {"product": "gamma", "price": "9", "quantity": "9"}]Type coercion at load time is one of the boring jobs pandas does for you. We'll meet it in Pandas.
5. Loading JSON
JSON maps directly to Python dict / list / scalars. No conversion needed.
import json json_text = ''' { "store": "Camden", "date": "2026-05-12", "items": [ {"sku": "C1", "qty": 4}, {"sku": "T2", "qty": 7} ] } ''' data = json.loads(json_text) # str → Python object print(data["store"]) # Camden print(data["items"][0]["qty"]) # 4 # Write it back out, pretty-printed print(json.dumps(data, indent=2))
For files, use json.load(f) / json.dump(obj, f). Path handling — see File I/O — pairs naturally with both formats.
6. The Mental Shift: Row-by-Row → Vectorised
Below is the most important slide in this lesson. Same task, two mindsets:
# Row-by-row — natural in Python, slow at scale prices = [3.50, 2.75, 4.00, 9.99, 14.50] with_vat = [] for p in prices: with_vat.append(p * 1.2) # Comprehension — same idea, more concise with_vat = [p * 1.2 for p in prices] # Vectorised — what NumPy will let you write next lesson import numpy as np arr = np.array(prices) with_vat = arr * 1.2 # one operation on the whole array
The first two iterate in Python (slow). The third pushes the loop into compiled C (fast). The mental shift — "I am operating on a whole column, not on each row" — is what separates DS code from general Python.
7. Mini-Example A: Word Count from Strings
from collections import Counter reviews = [ "great coffee fast service", "coffee was cold service slow", "great location great coffee", ] words = [w for review in reviews for w in review.split()] top = Counter(words).most_common(3) print(top) # [('coffee', 3), ('great', 3), ('service', 2)]
Nested comprehension flattens; Counter tallies; most_common ranks. Three primitives, one insight.
8. Mini-Example B: Per-Category Averages
import csv from io import StringIO from collections import defaultdict from statistics import mean csv_text = """category,price books,12.50 books,8.99 games,49.99 games,29.99 books,15.00 games,39.99 """ reader = csv.DictReader(StringIO(csv_text)) by_cat = defaultdict(list) for row in reader: by_cat[row["category"]].append(float(row["price"])) avg_by_cat = {cat: round(mean(prices), 2) for cat, prices in by_cat.items()} print(avg_by_cat) # {'books': 12.16, 'games': 39.99}
That's groupby-then-mean written by hand. Pandas will reduce it to one line — but doing it manually once cements what the abstraction actually does.
9. Generators: Streaming Big Files
A list holds every row in memory. A generator yields one at a time. For a 5 GB CSV, that's the difference between "fits" and "doesn't".
def stream_prices(rows): """Yield (sku, price) pairs without building a list.""" for row in rows: yield row["sku"], float(row["price"]) # Iterate without ever materialising the full list # for sku, price in stream_prices(reader): # ...
The pattern: build a pipeline of generators, then sum, max, or min over the final stage. Memory stays flat regardless of file size.
10. Jupyter Notebooks — A One-Paragraph Intro
A Jupyter notebook is a .ipynb file containing alternating cells of code and markdown. You run cells in any order; each cell's output (numbers, charts, tables) displays inline. Notebooks dominate exploratory DS because they let you think out loud — code, plot, prose, code, plot, prose — without rebuilding state. For shipped pipelines, you graduate back to .py files. Use both: notebooks for exploration, scripts for production.
Common Mistakes
- Writing
for i in range(len(x))when you meantenumerate(x). Always reach forenumerateorzipfirst. - Nested for-loops that a comprehension would handle. Two levels deep is the limit before readability suffers — three or more, write a real loop.
- Loading a giant CSV into a list. If you only need a sum or a max, stream it with a generator.
- Forgetting CSV values are strings.
row["quantity"] > 10is comparing a string to an int — always coerce types afterDictReader. - Reinventing
Counter. If you find yourself writingd[k] = d.get(k, 0) + 1, stop and importCounter.
🎯 Your Turn — top_categories
Write a function that takes CSV-like rows and returns the top N categories by total value. Standard library only.
Skeleton:
from collections import defaultdict def top_categories(rows, key_col, value_col, n=5): """Return [(category, total), ...] sorted desc, top N.""" # TODO 1: accumulate totals per category # TODO 2: sort by total descending and return the top n ... # Test rows = [ {"category": "books", "amount": "12.50"}, {"category": "games", "amount": "49.99"}, {"category": "books", "amount": "8.99"}, {"category": "music", "amount": "5.00"}, {"category": "games", "amount": "29.99"}, {"category": "books", "amount": "15.00"}, ] print(top_categories(rows, "category", "amount", n=2)) # Expected: [('games', 79.98), ('books', 36.49)]
Hint 1 — Accumulating
defaultdict(float) lets you write totals[cat] += float(row[value_col]) without checking if cat exists yet. Or use Counter — it sums floats too.
Hint 2 — Sorting
sorted(d.items(), key=lambda kv: kv[1], reverse=True)[:n] sorts a dict's items by value descending and slices the top n. Round the totals if you want clean output.
Show full solution
from collections import defaultdict def top_categories(rows, key_col, value_col, n=5): totals = defaultdict(float) for row in rows: totals[row[key_col]] += float(row[value_col]) ranked = sorted(totals.items(), key=lambda kv: kv[1], reverse=True) return [(cat, round(total, 2)) for cat, total in ranked[:n]] rows = [ {"category": "books", "amount": "12.50"}, {"category": "games", "amount": "49.99"}, {"category": "books", "amount": "8.99"}, {"category": "music", "amount": "5.00"}, {"category": "games", "amount": "29.99"}, {"category": "books", "amount": "15.00"}, ] print(top_categories(rows, "category", "amount", n=2)) # [('games', 79.98), ('books', 36.49)]
That's a groupby → sum → sort → head pipeline written four ways at once. In pandas, the same operation is one line: df.groupby("category")["amount"].sum().nlargest(n) — which you'll write next lesson but ahead of next.
What You Learned
- Comprehensions,
enumerate,zip,Countercover most preprocessing before pandas. csv.DictReaderandjson.loadsload the two formats you'll meet every day.- CSV values arrive as strings — type coercion is your job.
- Vectorised thinking — one operation on the whole column — is what NumPy unlocks next.
- Generators keep memory flat for huge files.
- Jupyter is for exploration; scripts are for production.
Next: NumPy: Working with Numbers — where the for-loop disappears and arrays do the work.