Your Data, Your Model
1 · The lesson
readEvery supervised ML problem boils down to one contract: given X, predict y.
X is what the model sees. y is what you want it to predict. Get this shape right and 80% of scikit-learn just works. Get it wrong and you'll spend an afternoon wondering why a "shape mismatch" error keeps shouting at you.
This lesson is the bridge between "I have a CSV" and "I have a model that predicts something". The minimum viable ML script lives in this lesson.
Run these right here — scikit-learn, pandas and matplotlib all work in the browser. The first scikit-learn import takes a few seconds while it downloads; after that it's instant. Expected output is also shown in comments below each block.
1. Features and Target
The convention every ML library on Earth follows:
X = features (the inputs, plural) y = target (the thing you predict, singular)
Imagine predicting house prices:
| sqft | bedrooms | age | price |
|---|---|---|---|
| 1200 | 2 | 12 | 320,000 |
| 1800 | 3 | 5 | 480,000 |
| 950 | 1 | 30 | 210,000 |
Xis the first three columns:sqft,bedrooms,age. Shape(3, 3)— three rows, three features.yis the last column:price. Shape(3,)— one value per row.
In NumPy / pandas terms:
import pandas as pd df = pd.DataFrame({ "sqft": [1200, 1800, 950], "bedrooms": [2, 3, 1], "age": [12, 5, 30], "price": [320_000, 480_000, 210_000], }) X = df[["sqft", "bedrooms", "age"]] y = df["price"] print(X.shape, y.shape) # → (3, 3) (3,)
The capital-X-lowercase-y convention is a hold-over from maths notation — X is a matrix, y is a vector. Every example you read uses it. Use it too.
2. The Train/Test Split
A model that "predicts" what it has already seen is useless — that's just a lookup table. The whole point is generalisation: doing well on data it has never seen.
So before training, you carve off a chunk of data and hide it. Train on the rest. Score on the hidden chunk. That score is your honest estimate of real-world performance.
from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 )
setup added so this can run · defines X, y
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X = _AutoMock('X') y = _AutoMock('y')
test_size=0.2— keep 20% for testing.random_state=42— fix the shuffle so the split is reproducible. Use any integer; pick one and stick with it across the project.
After this line you have four arrays:
X_train, y_train → model sees these during .fit()
X_test, y_test → model never sees these until evaluationIf you train on test data, your score is meaningless. This is the cardinal sin of ML — don't grade your own homework.
3. The sklearn API in Three Methods
Every scikit-learn estimator — and there are hundreds — has the same three methods:
model.fit(X_train, y_train) # learn from data model.predict(X_test) # make predictions on new data model.score(X_test, y_test) # one-number performance metric
setup added so this can run · defines X_train, y_train, X_test, y_test, model
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_train = _AutoMock('X_train') y_train = _AutoMock('y_train') X_test = _AutoMock('X_test') y_test = _AutoMock('y_test') model = _AutoMock('model')
That's the entire API contract. Once you've written this for one model, you've written it for all of them. Swap LinearRegression() for RandomForestRegressor() and the surrounding code is identical.
.score() returns a sensible default metric:
- For classifiers, it's accuracy.
- For regressors, it's R² (variance explained, 1.0 is perfect, 0.0 means "no better than predicting the mean").
You'll often want better metrics than .score() — we cover those in lesson 4 — but it's a fine first pass.
4. The Minimum Viable ML Script
Ten lines. Memorise the shape.
from sklearn.datasets import load_iris from sklearn.model_selection import train_test_split from sklearn.linear_model import LogisticRegression X, y = load_iris(return_X_y=True) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) model = LogisticRegression(max_iter=1000) model.fit(X_train, y_train) print(f"Test accuracy: {model.score(X_test, y_test):.2f}") # → Test accuracy: 1.00
Look at the structure:
1. Load the data.
2. Split it.
3. Instantiate the model.
4. Fit on the training data.
5. Score on the test data.
This shape works for 95% of classical ML problems. You'll change the model, the data, the metric — the skeleton stays the same.
5. Underfit, Overfit, and the Goldilocks Zone
When a model performs badly, it's failing in one of two ways.
Underfitting — too simple
The model is so weak it can't even capture the training data. Training error is high, test error is high. Imagine drawing a straight line through a clearly curved cloud of points.
Overfitting — too complex
The model memorises the training data, including the noise. Training error is tiny, test error is huge. Imagine drawing a wiggly line that passes through every single training point — perfectly fitted to the past, useless for the future.
The classic visual: fitting a noisy sine wave with polynomials of different degrees.
import numpy as np import matplotlib.pyplot as plt from sklearn.linear_model import LinearRegression from sklearn.preprocessing import PolynomialFeatures from sklearn.pipeline import make_pipeline rng = np.random.default_rng(42) X = np.linspace(0, 2 * np.pi, 30).reshape(-1, 1) y = np.sin(X).ravel() + rng.normal(0, 0.2, size=30) for degree, label in [(1, "underfit"), (4, "just right"), (15, "overfit")]: model = make_pipeline(PolynomialFeatures(degree), LinearRegression()) model.fit(X, y) plt.scatter(X, y, alpha=0.5) grid = np.linspace(0, 2 * np.pi, 200).reshape(-1, 1) plt.plot(grid, model.predict(grid), label=f"deg {degree} — {label}") plt.legend(); plt.title("Underfit / Just right / Overfit"); plt.show()
You'd see:
- Degree 1 — a straight line through a curve. Underfit.
- Degree 4 — follows the sine smoothly. Just right.
- Degree 15 — wiggles through every point. Overfit.
The goldilocks zone is whichever complexity gives you the lowest test error. That's why we have a test set in the first place — to find it.
6. Validation Set vs Test Set
For real projects, two splits aren't enough. You want three roles:
- Train — fit the model.
- Validation — tune hyperparameters (the model's knobs).
- Test — final, untouched score. Only look at this once.
If you tune knobs using the test set, the test set becomes part of your training process and stops being honest. You've leaked.
The cheap version is two train_test_split calls:
X_temp, X_test, y_temp, y_test = train_test_split(X, y, test_size=0.2, random_state=42) X_train, X_val, y_train, y_val = train_test_split(X_temp, y_temp, test_size=0.25, random_state=42) # Now: 60% train, 20% validation, 20% test
setup added so this can run · defines train_test_split, X, y
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def train_test_split(*_a, **_kw): print('-> train_test_split() called') return _AutoMock('train_test_split()') X = _AutoMock('X') y = _AutoMock('y')
Cross-validation — when data is scarce
If you only have 200 rows, carving 40% off for validation+test stings. K-fold cross-validation lets every row take a turn as validation.
from sklearn.model_selection import cross_val_score scores = cross_val_score(model, X_train, y_train, cv=5) print(f"CV mean: {scores.mean():.3f}, std: {scores.std():.3f}") # → CV mean: 0.967, std: 0.021
setup added so this can run · defines model, X_train, y_train
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) model = _AutoMock('model') X_train = _AutoMock('X_train') y_train = _AutoMock('y_train')
cv=5— split training data into 5 folds.- Train on 4, score on 1. Repeat 5 times so every fold gets a turn.
- Return 5 scores. The mean is your honest estimate; the std tells you how stable it is.
Cross-validation is the default move for any project with under ~10,000 rows. We use it heavily in pipelines.
Common Mistakes
- Training on the test set. Sometimes accidental: you call
train_test_splitafter fitting a scaler on all the data. The test set was already "seen" through the scaler's statistics. Lesson 6 fixes this withPipeline. - Predicting on un-preprocessed data. Whatever you did to
X_train(scaling, encoding, imputing) must happen toX_testwith the train-fitted transformer. Forgetting this is the most common production bug in ML. - Reporting the training score. "My model is 99% accurate!" — on what? If it's training accuracy, you've told us nothing about generalisation.
- Not setting
random_state. Your colleague reruns your notebook, gets a different split, gets a different number, files a bug. Pin the seed.
🎯 Your Turn — Train Your First Classifier
Fill in the gaps to:
1. Generate a synthetic 2-class dataset.
2. Split it 80/20.
3. Fit any sklearn classifier of your choice.
4. Return (train_accuracy, test_accuracy) as a tuple, both rounded to 3 decimals.
from sklearn.datasets import make_classification from sklearn.model_selection import train_test_split from sklearn.linear_model import LogisticRegression # or any other def train_and_score(): X, y = make_classification(n_samples=500, n_features=10, n_informative=5, random_state=42) # TODO 1: split X, y into train/test (80/20), random_state=42 # TODO 2: instantiate a classifier # TODO 3: fit on training data # TODO 4: compute train_acc and test_acc using model.score(...) # TODO 5: return (round(train_acc, 3), round(test_acc, 3)) pass print(train_and_score()) # → expected something like (0.918, 0.870)
Hint 1 — the split
train_test_split(X, y, test_size=0.2, random_state=42) returns four arrays in the order X_train, X_test, y_train, y_test.
Hint 2 — two .score() calls
model.score(X_train, y_train) gives training accuracy. model.score(X_test, y_test) gives test accuracy. Both return a float between 0 and 1.
Show full solution
from sklearn.datasets import make_classification from sklearn.model_selection import train_test_split from sklearn.linear_model import LogisticRegression def train_and_score(): X, y = make_classification(n_samples=500, n_features=10, n_informative=5, random_state=42) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) model = LogisticRegression(max_iter=1000) model.fit(X_train, y_train) train_acc = model.score(X_train, y_train) test_acc = model.score(X_test, y_test) return (round(train_acc, 3), round(test_acc, 3)) print(train_and_score()) # → (0.918, 0.870)
Training accuracy slightly higher than test is normal and healthy. If they were identical, your model might be too simple. If training was 0.99 and test was 0.65, you'd be overfitting. The small gap here is the sweet spot.
What You Learned
- The contract:
X(features) andy(target). Capital X, lowercase y. Always. train_test_splitcarves off a hold-out chunk you never touch during training.- Every sklearn estimator has
.fit(),.predict(),.score(). Same shape, different models. - The minimum viable ML script: load → split → fit → score. Five steps.
- Underfit = too simple. Overfit = memorised noise. Goldilocks = best test score.
- Use a validation set (or cross-validation) for tuning, keep the test set untouched until the end.
Next: Linear Regression — predicting numbers with the simplest model in the toolkit.