Linear & Logistic Regression
1 · The lesson
readTwo tasks cover the vast majority of supervised learning in production: predict a number (regression) or predict a category (classification). Linear regression and logistic regression are the canonical models for each, and they remain the right starting point for nearly every tabular ML problem you'll meet — not because they always win, but because they set the floor every fancier model must clear.
This lesson walks the full pipeline — split, fit, evaluate, regularise, tune — for both regression and classification, with the metrics that matter and the traps that ambush juniors. By the end you'll know why accuracy lies, why scaling is non-negotiable for Ridge, and where data leakage hides.
Runtime note: scikit-learn, pandas and NumPy run right here — press Run on any block. The first scikit-learn import takes a few seconds while the package downloads; after that it's instant. Expected output is also shown in comments.
1. The Two Canonical Tasks
| Task | Output | Example | Loss |
|---|---|---|---|
| Regression | Real number | House price, daily revenue, temperature | MSE / MAE |
| Classification | Discrete label | Spam vs ham, churn vs stay, 10-class digit | Log-loss / cross-entropy |
Both fit the same shape — a weighted sum of inputs — and differ only in what they do at the end. Regression returns the sum directly. Classification squashes the sum through a non-linearity to get a probability, then thresholds it.
2. Linear Regression — Fit the Best Line
Given points (x, y), find the line y = w·x + b that minimises the sum of squared errors between predicted and actual y. With multiple inputs the line becomes a hyperplane:
y = w₁·x₁ + w₂·x₂ + ... + wₙ·xₙ + b
The weights wᵢ say how much y moves when xᵢ moves by one unit, all else held constant. The intercept b is the prediction when every xᵢ = 0.
# Local: pip install scikit-learn numpy import numpy as np from sklearn.linear_model import LinearRegression X = np.array([[1], [2], [3], [4], [5]]) # one feature, five samples y = np.array([2.1, 4.0, 6.1, 7.9, 10.2]) # roughly y = 2x model = LinearRegression() model.fit(X, y) print(model.coef_) # [2.03] — slope ~2 print(model.intercept_) # 0.01 — intercept ~0 print(model.predict([[6]])) # [12.19] — extrapolation print(model.score(X, y)) # 0.9994 — R²
fit(X, y) solves the normal equations under the hood — closed-form, no iteration. predict applies the learned weights. score returns R² (coefficient of determination) by default.
3. Multiple Features and Coefficient Interpretation
import numpy as np from sklearn.linear_model import LinearRegression # rows: samples; cols: [rooms, age_years, sqft] X = np.array([[3, 10, 1500], [4, 25, 2200], [2, 40, 900], [5, 5, 2800], [3, 15, 1700]]) y = np.array([420_000, 530_000, 220_000, 740_000, 460_000]) model = LinearRegression().fit(X, y) print(dict(zip(["rooms", "age", "sqft"], model.coef_))) # {'rooms': 12_530.4, 'age': -2_870.1, 'sqft': 198.7}
Read coefficients carefully: "holding age and sqft fixed, one extra room is worth ~£12.5k". Two pitfalls:
- Scale dependence —
sqfthas tiny coefficient but huge range;roomshas big coefficient but small range. You cannot rank features by raw coefficient magnitude unless you've standardised the inputs first (mean 0, std 1) — see feature. - Causation — a coefficient is a conditional association, not a causal effect. Adding a sunroom doesn't cause the price bump if your data merely correlates the two.
4. Train/Test Split — The Non-Negotiable
Evaluating a model on the data it learned from is like grading a student on the exam they wrote. You'll overestimate performance every time.
from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) model = LinearRegression().fit(X_train, y_train) print(model.score(X_test, y_test)) # honest R² on unseen data
setup added so this can run · defines X, y, LinearRegression
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X = _AutoMock('X') y = _AutoMock('y') def LinearRegression(*_a, **_kw): print('-> LinearRegression() called') return _AutoMock('LinearRegression()')
test_size=0.2— hold out 20% for evaluation.random_state=42— reproducible split. Pick any integer; use the same one across the team.- For time-series, never random-split — use
TimeSeriesSplitso the test set is always after training in time.
5. Regression Metrics — When to Use Which
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score import numpy as np y_pred = model.predict(X_test) mae = mean_absolute_error(y_test, y_pred) mse = mean_squared_error(y_test, y_pred) rmse = np.sqrt(mse) r2 = r2_score(y_test, y_pred)
setup added so this can run · defines X_test, y_test, model
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_test = _AutoMock('X_test') y_test = _AutoMock('y_test') model = _AutoMock('model')
| Metric | What it measures | Use when |
|---|---|---|
| MAE | Average absolute error, same units as y | You want robust, intuitive "off by £X on average" |
| MSE | Average squared error | Penalises big misses harder; the optimisation target |
| RMSE | √MSE, back in original units | Reporting; the de-facto industry metric |
| R² | Fraction of variance explained, in [-∞, 1] | Comparing to a "predict the mean" baseline |
R² > 0 means you beat the mean baseline; R² = 1 is perfect; R² can go negative if your model is worse than predicting the mean. RMSE > MAE always — if they diverge wildly, you have large outliers.
6. The Textbook Assumptions (and Reality)
Statistics textbooks insist linear regression assumes:
1. Linearity — the true relationship is linear in the inputs.
2. Independence — observations don't influence each other.
3. Homoscedasticity — error variance is constant across x.
4. Normality of residuals — errors are Gaussian.
5. No multicollinearity — features aren't redundant.
In practice, no production team checks all five. We check residual plots for obvious structure (curvature → add polynomial features; funnel shape → log-transform y), eyeball collinearity via correlation matrices or VIF, and rely on cross-validated test-set RMSE to tell us whether the model is useful. Worry about the assumptions when inference matters (p-values, confidence intervals); for pure prediction, generalisation error is king.
7. Regularisation — Ridge, Lasso, ElasticNet
When you have many features (or few samples), plain linear regression overfits — it memorises noise. Regularisation adds a penalty on coefficient size, shrinking weights toward zero.
from sklearn.linear_model import Ridge, Lasso, ElasticNet from sklearn.preprocessing import StandardScaler from sklearn.pipeline import Pipeline # Regularised models REQUIRE scaled features — coefficient magnitude is now meaningful ridge = Pipeline([ ("scale", StandardScaler()), ("model", Ridge(alpha=1.0)), # L2 — shrinks all coefs smoothly ]) lasso = Pipeline([ ("scale", StandardScaler()), ("model", Lasso(alpha=0.1)), # L1 — drives some coefs to ZERO (feature selection) ]) enet = Pipeline([ ("scale", StandardScaler()), ("model", ElasticNet(alpha=0.1, l1_ratio=0.5)), # mix of L1 and L2 ])
| Model | Penalty | Effect | When |
|---|---|---|---|
| Ridge | α · Σ wᵢ² | Shrinks coefs smoothly, keeps all features | Many correlated features; you want stability |
| Lasso | α · Σ \|wᵢ\| | Drives weak coefs to exactly 0 | You want automatic feature selection |
| ElasticNet | mix of both | Best of both | Correlated features + selection wanted |
alpha is the regularisation strength — bigger means more shrinkage. Tune it via cross-validation (Section 13).
Why scaling is mandatory: the penalty Σ wᵢ² punishes large weights. A feature on scale 0–1,000,000 will need a tiny weight to make any contribution and won't get penalised; a feature on scale 0–1 will need a huge weight and get crushed. Without scaling, regularisation is biased toward whichever features happen to have small numeric ranges.
8. Logistic Regression — Same Idea, Different Output
For binary classification (e.g. spam vs ham), compute the same linear combination and squash through the sigmoid:
z = w₁·x₁ + ... + wₙ·xₙ + b p = 1 / (1 + e⁻ᶻ) # in (0, 1) — interpret as P(class = 1)
Predict class 1 if p ≥ 0.5, else class 0. The 0.5 boundary is convention, not law (more on threshold tuning below).
from sklearn.linear_model import LogisticRegression from sklearn.datasets import make_classification from sklearn.model_selection import train_test_split X, y = make_classification(n_samples=1000, n_features=10, random_state=42) X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) clf = LogisticRegression(max_iter=1000).fit(X_train, y_train) print(clf.predict(X_test[:5])) # [0 1 1 0 1] print(clf.predict_proba(X_test[:5])) # probabilities per class print(clf.score(X_test, y_test)) # accuracy
Coefficients of logistic regression are log-odds: a one-unit increase in xᵢ multiplies the odds of class 1 by exp(wᵢ). Useful for inference; almost nobody bothers in pure prediction work.
9. Threshold Tuning
The default 0.5 threshold is fine when classes are balanced and false positives/negatives cost the same. They almost never do.
import numpy as np probs = clf.predict_proba(X_test)[:, 1] # P(class = 1) # Aggressive — flag anything with > 30% probability (high recall, more false positives) y_pred_aggressive = (probs > 0.30).astype(int) # Conservative — only flag if > 70% confident (high precision, more false negatives) y_pred_conservative = (probs > 0.70).astype(int)
setup added so this can run · defines X_test, clf
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_test = _AutoMock('X_test') clf = _AutoMock('clf')
Cancer screening tolerates false positives (re-test); fraud blocking tolerates false negatives (manual review queue) — pick the threshold by cost, not by 0.5 reflex.
10. Classification Metrics — Why Accuracy Lies
from sklearn.metrics import ( accuracy_score, precision_score, recall_score, f1_score, roc_auc_score, confusion_matrix, classification_report ) y_pred = clf.predict(X_test) print(confusion_matrix(y_test, y_pred)) # [[TN FP] # [FN TP]] print(classification_report(y_test, y_pred)) # precision recall f1-score support # 0 0.91 0.89 0.90 100 # 1 0.89 0.91 0.90 100 # accuracy 0.90 200
setup added so this can run · defines X_test, clf, y_test
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_test = _AutoMock('X_test') clf = _AutoMock('clf') y_test = _AutoMock('y_test')
| Metric | Formula | Use when |
|---|---|---|
| Accuracy | (TP+TN) / total | Balanced classes, equal costs |
| Precision | TP / (TP+FP) | False positives are expensive (spam filter — don't flag real mail) |
| Recall | TP / (TP+FN) | False negatives are expensive (cancer — don't miss a case) |
| F1 | harmonic mean of P and R | Balance both; default for imbalanced |
| ROC-AUC | area under ROC curve | Ranking quality across all thresholds |
Accuracy lies on imbalanced data. If 99% of transactions are legitimate, a model that predicts "legitimate" for everything scores 99% accuracy and catches zero fraud. Always check the confusion matrix; for imbalanced problems, optimise F1 or AUC (or precision/recall at a chosen threshold tied to business cost).
11. ROC-AUC in One Snippet
from sklearn.metrics import roc_auc_score, roc_curve probs = clf.predict_proba(X_test)[:, 1] auc = roc_auc_score(y_test, probs) print(auc) # 0.95 — closer to 1 is better; 0.5 is random fpr, tpr, thresholds = roc_curve(y_test, probs)
setup added so this can run · defines y_test, X_test, clf
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) y_test = _AutoMock('y_test') X_test = _AutoMock('X_test') clf = _AutoMock('clf')
AUC = 0.5 is coin-flip; 1.0 is perfect. It measures the probability that a random positive is ranked higher than a random negative — threshold-free, so it's the right summary when you haven't picked a threshold yet.
12. Cross-Validation — One Split Isn't Enough
A single train/test split is one sample of model performance. The variance across splits can be large — your "85% accuracy" might be 78–92% depending on the seed. k-fold cross-validation runs k train/test splits across the dataset and averages the scores.
from sklearn.model_selection import cross_val_score scores = cross_val_score(LogisticRegression(max_iter=1000), X, y, cv=5, scoring="f1") print(scores) # [0.89 0.91 0.87 0.90 0.88] print(scores.mean(), scores.std()) # 0.89 ± 0.014
setup added so this can run · defines X, y, LogisticRegression
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X = _AutoMock('X') y = _AutoMock('y') def LogisticRegression(*_a, **_kw): print('-> LogisticRegression() called') return _AutoMock('LogisticRegression()')
cv=5 is the default sweet spot. Use StratifiedKFold (sklearn's default for classifiers) so each fold preserves the class ratio. For time-series, TimeSeriesSplit. Report mean ± std — a single number hides the variance.
13. Hyperparameter Tuning with GridSearchCV
from sklearn.model_selection import GridSearchCV from sklearn.linear_model import Ridge from sklearn.preprocessing import StandardScaler from sklearn.pipeline import Pipeline pipe = Pipeline([("scale", StandardScaler()), ("model", Ridge())]) param_grid = {"model__alpha": [0.001, 0.01, 0.1, 1, 10, 100]} search = GridSearchCV(pipe, param_grid, cv=5, scoring="neg_root_mean_squared_error") search.fit(X_train, y_train) print(search.best_params_) # {'model__alpha': 1.0} print(-search.best_score_) # 12_340.5 (RMSE) print(search.score(X_test, y_test)) # held-out test score
setup added so this can run · defines X_train, y_train, X_test, y_test
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_train = _AutoMock('X_train') y_train = _AutoMock('y_train') X_test = _AutoMock('X_test') y_test = _AutoMock('y_test')
Key points:
model__alpha— the double underscore selects thealphaparameter on the step namedmodelinside the pipeline.- Scoring is negated (
neg_root_mean_squared_error) because GridSearchCV always maximises — for losses you want minimised, flip the sign. - Always fit the pipeline inside the CV loop — scaling-before-split leaks test statistics into training.
For large grids, RandomizedSearchCV samples instead of exhaustively trying every combination.
14. Multi-Class Logistic
Three or more classes? Two strategies:
LogisticRegression(multi_class="ovr") # one-vs-rest — train K binary classifiers LogisticRegression(multi_class="multinomial") # true softmax — joint probability over K classes
setup added so this can run · defines LogisticRegression
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def LogisticRegression(*_a, **_kw): print('-> LogisticRegression() called') return _AutoMock('LogisticRegression()')
multinomial (softmax) is usually more accurate and well-calibrated; ovr is faster to train and easier to explain. Sklearn picks multinomial automatically with the lbfgs solver for most datasets.
Common Mistakes
1. Not scaling features for regularised regression. Ridge/Lasso/ElasticNet penalties depend on coefficient magnitude. Unscaled features → biased penalty → wrong coefs. Always pipeline StandardScaler before the model. Plain LinearRegression doesn't care.
2. Accuracy on imbalanced data.
# 99% class 0, 1% class 1 — predict-everything-zero y_pred = np.zeros_like(y_test) print(accuracy_score(y_test, y_pred)) # 0.99 — meaningless print(recall_score(y_test, y_pred)) # 0 — caught zero positives
setup added so this can run · defines y_test, np, accuracy_score, recall_score
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) y_test = _AutoMock('y_test') np = _AutoMock('np') def accuracy_score(*_a, **_kw): print('-> accuracy_score() called') return _AutoMock('accuracy_score()') def recall_score(*_a, **_kw): print('-> recall_score() called') return _AutoMock('recall_score()')
Use F1, AUC, or recall@precision-floor instead. Set class_weight="balanced" to make the loss attend to the minority class.
3. Predicting on un-transformed data. If you scaled X_train with StandardScaler, you must apply the same fitted scaler to X_test (and any new prediction input). Easiest fix: wrap everything in a Pipeline — fit-transform on train, transform-only on test happens automatically.
4. Data leakage via pre-split transformation.
# WRONG — scaler sees test statistics scaler = StandardScaler().fit(X) # fit on EVERYTHING X_scaled = scaler.transform(X) X_train, X_test, y_train, y_test = train_test_split(X_scaled, y, ...)
setup added so this can run · defines X, train_test_split, y, StandardScaler
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X = _AutoMock('X') def train_test_split(*_a, **_kw): print('-> train_test_split() called') return _AutoMock('train_test_split()') y = _AutoMock('y') def StandardScaler(*_a, **_kw): print('-> StandardScaler() called') return _AutoMock('StandardScaler()')
The scaler has now memorised the test set's mean and std — your reported test score is optimistic. Always split first, then fit transformers on X_train only. Pipelines + cross-validation do this correctly out of the box.
5. Coefficient hunting. Reading raw model.coef_ without (a) standardising features or (b) statistical significance tests is a beginner trap. For ranking feature importance honestly, use permutation importance (covered in trees).
6. max_iter warnings. LogisticRegression defaults to 100 iterations; large/unscaled problems hit the cap before converging — sklearn warns. Either scale the data or bump to max_iter=1000.
🎯 Your Turn — A Full Ridge Pipeline
Build a complete regression workflow on the inline boston_lite data: split, scale, fit Ridge with α tuned via 5-fold CV, predict on the held-out set, report RMSE and R².
Requirements:
1. Use train_test_split with test_size=0.2, random_state=42.
2. Wrap StandardScaler and Ridge in a Pipeline.
3. Use GridSearchCV to tune alpha over [0.01, 0.1, 1, 10, 100] with 5-fold CV.
4. Report best_params_, the test-set RMSE, and the test-set R².
Skeleton:
import numpy as np from sklearn.linear_model import Ridge from sklearn.preprocessing import StandardScaler from sklearn.pipeline import Pipeline from sklearn.model_selection import train_test_split, GridSearchCV from sklearn.metrics import mean_squared_error, r2_score # Inline mini dataset: [rooms, age, lstat] -> median house value (in $1000s) X = np.array([ [6.575, 65.2, 4.98], [6.421, 78.9, 9.14], [7.185, 61.1, 4.03], [6.998, 45.8, 2.94], [7.147, 54.2, 5.33], [6.430, 58.7, 5.21], [6.012, 66.6, 12.43], [6.172, 96.1, 19.15], [5.631, 100.0, 29.93], [6.004, 85.9, 17.10], [6.377, 70.3, 20.45], [6.009, 94.3, 13.27], [5.889, 39.0, 15.71], [7.249, 21.9, 4.81], [6.625, 17.5, 3.81], [5.927, 42.6, 18.66], [6.137, 32.1, 13.44], [5.951, 38.5, 12.13], [6.456, 71.9, 11.10], [6.815, 40.5, 5.25], ]) y = np.array([24.0, 21.6, 34.7, 33.4, 36.2, 28.7, 22.9, 17.5, 16.5, 18.9, 15.0, 18.9, 21.7, 35.4, 28.7, 19.6, 23.1, 21.7, 22.2, 30.8]) # TODO 1: split into train/test (test_size=0.2, random_state=42) # TODO 2: build Pipeline([("scale", StandardScaler()), ("model", Ridge())]) # TODO 3: GridSearchCV over model__alpha with cv=5, scoring="neg_root_mean_squared_error" # TODO 4: fit on training data, predict on test # TODO 5: print best_params_, RMSE, R²
Hint 1 — Pipeline parameter names
InsideGridSearchCV, parameters belonging to a pipeline step are named stepname__paramname — two underscores. So Ridge's alpha inside a step called "model" becomes "model__alpha". Get this wrong and you'll see a baffling ValueError: Invalid parameter.
Hint 2 — RMSE from MSE
mean_squared_error returns MSE by default. Either pass squared=False (sklearn ≥ 0.22) to get RMSE directly, or take np.sqrt(mean_squared_error(...)). Also remember GridSearchCV's best_score_ is negated — flip its sign before reporting.
Show full solution
import numpy as np from sklearn.linear_model import Ridge from sklearn.preprocessing import StandardScaler from sklearn.pipeline import Pipeline from sklearn.model_selection import train_test_split, GridSearchCV from sklearn.metrics import mean_squared_error, r2_score X = np.array([ [6.575, 65.2, 4.98], [6.421, 78.9, 9.14], [7.185, 61.1, 4.03], [6.998, 45.8, 2.94], [7.147, 54.2, 5.33], [6.430, 58.7, 5.21], [6.012, 66.6, 12.43], [6.172, 96.1, 19.15], [5.631, 100.0, 29.93], [6.004, 85.9, 17.10], [6.377, 70.3, 20.45], [6.009, 94.3, 13.27], [5.889, 39.0, 15.71], [7.249, 21.9, 4.81], [6.625, 17.5, 3.81], [5.927, 42.6, 18.66], [6.137, 32.1, 13.44], [5.951, 38.5, 12.13], [6.456, 71.9, 11.10], [6.815, 40.5, 5.25], ]) y = np.array([24.0, 21.6, 34.7, 33.4, 36.2, 28.7, 22.9, 17.5, 16.5, 18.9, 15.0, 18.9, 21.7, 35.4, 28.7, 19.6, 23.1, 21.7, 22.2, 30.8]) X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) pipe = Pipeline([ ("scale", StandardScaler()), ("model", Ridge()), ]) param_grid = {"model__alpha": [0.01, 0.1, 1.0, 10.0, 100.0]} search = GridSearchCV( pipe, param_grid, cv=5, scoring="neg_root_mean_squared_error", ) search.fit(X_train, y_train) y_pred = search.predict(X_test) rmse = np.sqrt(mean_squared_error(y_test, y_pred)) r2 = r2_score(y_test, y_pred) print(f"Best alpha: {search.best_params_['model__alpha']}") print(f"CV RMSE: {-search.best_score_:.3f}") print(f"Test RMSE: {rmse:.3f}") print(f"Test R²: {r2:.3f}") # Expected output (approximate — small dataset, results vary slightly): # Best alpha: 1.0 # CV RMSE: 4.1 # Test RMSE: 3.8 # Test R²: 0.72
Why this solution is right by construction:
- Split first — the scaler never sees
X_test, no leakage. - Pipeline — the same scaler that fit on
X_trainis automatically applied at predict time onX_test. No risk of forgetting to transform. - Cross-validated tuning —
alpha=1.0is picked because it gave the best mean RMSE across the 5 folds of the training set, not because of one lucky split. - Final evaluation on held-out test — the only honest performance number is the one on data that touched neither the scaler's
.fit()nor the alpha search.
Drop in Lasso instead of Ridge and you get automatic feature selection — coefficients of weak features become exactly zero. Same pipeline, one-word change.
What You Learned
- Regression predicts a number; classification predicts a category. Linear and logistic regression are the canonical starting models for each.
- Linear regression fits
y = w·x + bby minimising squared errors; coefficients are interpretable only after standardising features. train_test_splitwith a fixedrandom_stategives you an honest, reproducible held-out evaluation.- Regression metrics: RMSE for industry reporting, MAE for robustness to outliers, R² to compare against the mean baseline.
- Regularisation — Ridge (L2) shrinks smoothly; Lasso (L1) zeroes out weak features; ElasticNet mixes both. All three require scaled inputs.
- Logistic regression = linear combination → sigmoid → probability → threshold. Default 0.5 is convention, not law.
- Classification metrics: accuracy lies on imbalanced data — use F1, AUC, or precision/recall tied to business cost.
- k-fold cross-validation beats a single split for estimating performance and tuning hyperparameters.
GridSearchCV+Pipelineis the safe, leakage-free pattern for hyperparameter tuning.- Multi-class logistic uses one-vs-rest or softmax (multinomial); sklearn picks sensibly by default.
Next: Decision Trees & Random Forests — non-linear models, ensembles, and the trees that beat linear baselines on most tabular data.