PythonMastery
beginner 16 min read · lesson 2 of 6 in Data Science Fundamentals

Python for Data Analysis

1 · The lesson

read

Before NumPy and pandas arrive, Python's standard library already does a surprising amount of analysis work. This lesson covers the DS-flavoured idioms — comprehensions, zip, Counter, generators — and the file formats you'll meet on day one. By the end you'll have written a tiny analytics script with zero third-party imports.


1. Comprehensions Are the Lingua Franca

A data analysis script transforms one collection into another. Comprehensions express that in one line. You met them in Comprehensions; here's how DS code uses them.

python
prices = [9.99, 14.50, 7.25, 22.00, 3.99]

# Transform — add VAT to each price
with_vat = [round(p * 1.2, 2) for p in prices]
print(with_vat)                      # [11.99, 17.4, 8.7, 26.4, 4.79]

# Filter — only items above £10
expensive = [p for p in prices if p > 10]
print(expensive)                     # [14.5, 22.0]

# Map to a dict
sku_to_price = {f"SKU-{i:03d}": p for i, p in enumerate(prices, start=1)}
print(sku_to_price)
# {'SKU-001': 9.99, 'SKU-002': 14.5, 'SKU-003': 7.25, 'SKU-004': 22.0, 'SKU-005': 3.99}

# Aggregate — total revenue if we sold one of each
total = sum(p for p in prices)
print(f"Total: £{total:.2f}")        # Total: £57.73

Memorise the shapes: [expr for x in iter if cond], {k: v for ...}, sum(expr for ...). They cover most cleanup tasks before pandas enters the room.


2. enumerate and zip — The Pair You Actually Use

enumerate gives you (index, value) as you loop. zip walks several iterables in lockstep. Together they replace 80% of range(len(...))-style code.

python
products = ["Coffee", "Tea", "Cocoa"]
prices   = [3.50, 2.75, 4.00]
units    = [120,  340,   85]

# enumerate — numbered output
for i, name in enumerate(products, start=1):
    print(f"{i}. {name}")

# zip — pair columns together
revenue = {name: price * qty for name, price, qty in zip(products, prices, units)}
print(revenue)
# {'Coffee': 420.0, 'Tea': 935.0, 'Cocoa': 340.0}

zip stops at the shortest iterable — pass strict=True (Python 3.10+) if you want it to raise when lengths mismatch.


3. collections.Counter — Tallies in One Line

Whenever you'd write if key in d: d[key] += 1 else: d[key] = 1, reach for Counter.

python
from collections import Counter

orders = ["coffee", "tea", "coffee", "coffee", "cocoa", "tea", "coffee"]

tally = Counter(orders)
print(tally)                         # Counter({'coffee': 4, 'tea': 2, 'cocoa': 1})
print(tally.most_common(2))          # [('coffee', 4), ('tea', 2)]
print(tally["coffee"])               # 4
print(tally["water"])                # 0  — missing keys return 0, never raise

most_common() is the one-line "top-N categories" function you'll write a thousand times. Counter also supports arithmetic — Counter(a) + Counter(b) merges two tallies.


4. Loading CSV with the Standard Library

Most DS data arrives as CSV. The stdlib csv module handles it without any dependency. (Recap and details in CSV & JSON.)

python
import csv
from io import StringIO

# A tiny CSV inline so we don't need a file
csv_text = """product,price,quantity
Coffee,3.50,120
Tea,2.75,340
Cocoa,4.00,85
"""

# DictReader yields each row as {column: value}
reader = csv.DictReader(StringIO(csv_text))
rows = list(reader)

for row in rows:
    print(row)
# {'product': 'Coffee', 'price': '3.50', 'quantity': '120'}
# ...

Note the values are strings — CSV has no types. You convert them yourself:

python
typed = [
    {"product": r["product"], "price": float(r["price"]), "quantity": int(r["quantity"])}
    for r in rows
]
+ setup added so this can run · defines rows
rows = [{"product": "alpha", "price": "3", "quantity": "3"}, {"product": "beta", "price": "6", "quantity": "6"}, {"product": "gamma", "price": "9", "quantity": "9"}]

Type coercion at load time is one of the boring jobs pandas does for you. We'll meet it in Pandas.


5. Loading JSON

JSON maps directly to Python dict / list / scalars. No conversion needed.

python
import json

json_text = '''
{
    "store": "Camden",
    "date":  "2026-05-12",
    "items": [
        {"sku": "C1", "qty": 4},
        {"sku": "T2", "qty": 7}
    ]
}
'''

data = json.loads(json_text)         # str → Python object
print(data["store"])                 # Camden
print(data["items"][0]["qty"])       # 4

# Write it back out, pretty-printed
print(json.dumps(data, indent=2))

For files, use json.load(f) / json.dump(obj, f). Path handling — see File I/O — pairs naturally with both formats.


6. The Mental Shift: Row-by-Row → Vectorised

Below is the most important slide in this lesson. Same task, two mindsets:

python
# Row-by-row — natural in Python, slow at scale
prices = [3.50, 2.75, 4.00, 9.99, 14.50]
with_vat = []
for p in prices:
    with_vat.append(p * 1.2)

# Comprehension — same idea, more concise
with_vat = [p * 1.2 for p in prices]

# Vectorised — what NumPy will let you write next lesson
import numpy as np
arr = np.array(prices)
with_vat = arr * 1.2                 # one operation on the whole array

The first two iterate in Python (slow). The third pushes the loop into compiled C (fast). The mental shift — "I am operating on a whole column, not on each row" — is what separates DS code from general Python.


7. Mini-Example A: Word Count from Strings

python
from collections import Counter

reviews = [
    "great coffee fast service",
    "coffee was cold service slow",
    "great location great coffee",
]

words = [w for review in reviews for w in review.split()]
top = Counter(words).most_common(3)
print(top)
# [('coffee', 3), ('great', 3), ('service', 2)]

Nested comprehension flattens; Counter tallies; most_common ranks. Three primitives, one insight.


8. Mini-Example B: Per-Category Averages

python
import csv
from io import StringIO
from collections import defaultdict
from statistics import mean

csv_text = """category,price
books,12.50
books,8.99
games,49.99
games,29.99
books,15.00
games,39.99
"""

reader = csv.DictReader(StringIO(csv_text))

by_cat = defaultdict(list)
for row in reader:
    by_cat[row["category"]].append(float(row["price"]))

avg_by_cat = {cat: round(mean(prices), 2) for cat, prices in by_cat.items()}
print(avg_by_cat)
# {'books': 12.16, 'games': 39.99}

That's groupby-then-mean written by hand. Pandas will reduce it to one line — but doing it manually once cements what the abstraction actually does.


9. Generators: Streaming Big Files

A list holds every row in memory. A generator yields one at a time. For a 5 GB CSV, that's the difference between "fits" and "doesn't".

python
def stream_prices(rows):
    """Yield (sku, price) pairs without building a list."""
    for row in rows:
        yield row["sku"], float(row["price"])

# Iterate without ever materialising the full list
# for sku, price in stream_prices(reader):
#     ...

The pattern: build a pipeline of generators, then sum, max, or min over the final stage. Memory stays flat regardless of file size.


10. Jupyter Notebooks — A One-Paragraph Intro

A Jupyter notebook is a .ipynb file containing alternating cells of code and markdown. You run cells in any order; each cell's output (numbers, charts, tables) displays inline. Notebooks dominate exploratory DS because they let you think out loud — code, plot, prose, code, plot, prose — without rebuilding state. For shipped pipelines, you graduate back to .py files. Use both: notebooks for exploration, scripts for production.


Common Mistakes

  • Writing for i in range(len(x)) when you meant enumerate(x). Always reach for enumerate or zip first.
  • Nested for-loops that a comprehension would handle. Two levels deep is the limit before readability suffers — three or more, write a real loop.
  • Loading a giant CSV into a list. If you only need a sum or a max, stream it with a generator.
  • Forgetting CSV values are strings. row["quantity"] > 10 is comparing a string to an int — always coerce types after DictReader.
  • Reinventing Counter. If you find yourself writing d[k] = d.get(k, 0) + 1, stop and import Counter.

🎯 Your Turn — top_categories

Write a function that takes CSV-like rows and returns the top N categories by total value. Standard library only.

Skeleton:

python
from collections import defaultdict

def top_categories(rows, key_col, value_col, n=5):
    """Return [(category, total), ...] sorted desc, top N."""
    # TODO 1: accumulate totals per category
    # TODO 2: sort by total descending and return the top n
    ...

# Test
rows = [
    {"category": "books", "amount": "12.50"},
    {"category": "games", "amount": "49.99"},
    {"category": "books", "amount": "8.99"},
    {"category": "music", "amount": "5.00"},
    {"category": "games", "amount": "29.99"},
    {"category": "books", "amount": "15.00"},
]
print(top_categories(rows, "category", "amount", n=2))
# Expected: [('games', 79.98), ('books', 36.49)]
Hint 1 — Accumulating defaultdict(float) lets you write totals[cat] += float(row[value_col]) without checking if cat exists yet. Or use Counter — it sums floats too.
Hint 2 — Sorting sorted(d.items(), key=lambda kv: kv[1], reverse=True)[:n] sorts a dict's items by value descending and slices the top n. Round the totals if you want clean output.
Show full solution
python
from collections import defaultdict

def top_categories(rows, key_col, value_col, n=5):
    totals = defaultdict(float)
    for row in rows:
        totals[row[key_col]] += float(row[value_col])
    ranked = sorted(totals.items(), key=lambda kv: kv[1], reverse=True)
    return [(cat, round(total, 2)) for cat, total in ranked[:n]]

rows = [
    {"category": "books", "amount": "12.50"},
    {"category": "games", "amount": "49.99"},
    {"category": "books", "amount": "8.99"},
    {"category": "music", "amount": "5.00"},
    {"category": "games", "amount": "29.99"},
    {"category": "books", "amount": "15.00"},
]
print(top_categories(rows, "category", "amount", n=2))
# [('games', 79.98), ('books', 36.49)]

That's a groupby → sum → sort → head pipeline written four ways at once. In pandas, the same operation is one line: df.groupby("category")["amount"].sum().nlargest(n) — which you'll write next lesson but ahead of next.


What You Learned

  • Comprehensions, enumerate, zip, Counter cover most preprocessing before pandas.
  • csv.DictReader and json.loads load the two formats you'll meet every day.
  • CSV values arrive as strings — type coercion is your job.
  • Vectorised thinking — one operation on the whole column — is what NumPy unlocks next.
  • Generators keep memory flat for huge files.
  • Jupyter is for exploration; scripts are for production.

Next: NumPy: Working with Numbers — where the for-loop disappears and arrays do the work.