PythonMastery
intermediate 22 min read · lesson 7 of 9 in Data Science & ML

Linear & Logistic Regression

1 · The lesson

read

Two tasks cover the vast majority of supervised learning in production: predict a number (regression) or predict a category (classification). Linear regression and logistic regression are the canonical models for each, and they remain the right starting point for nearly every tabular ML problem you'll meet — not because they always win, but because they set the floor every fancier model must clear.

This lesson walks the full pipeline — split, fit, evaluate, regularise, tune — for both regression and classification, with the metrics that matter and the traps that ambush juniors. By the end you'll know why accuracy lies, why scaling is non-negotiable for Ridge, and where data leakage hides.

Runtime note: scikit-learn, pandas and NumPy run right here — press Run on any block. The first scikit-learn import takes a few seconds while the package downloads; after that it's instant. Expected output is also shown in comments.


1. The Two Canonical Tasks

TaskOutputExampleLoss
RegressionReal numberHouse price, daily revenue, temperatureMSE / MAE
ClassificationDiscrete labelSpam vs ham, churn vs stay, 10-class digitLog-loss / cross-entropy

Both fit the same shape — a weighted sum of inputs — and differ only in what they do at the end. Regression returns the sum directly. Classification squashes the sum through a non-linearity to get a probability, then thresholds it.


2. Linear Regression — Fit the Best Line

Given points (x, y), find the line y = w·x + b that minimises the sum of squared errors between predicted and actual y. With multiple inputs the line becomes a hyperplane:

python
y = w₁·x₁ + w₂·x₂ + ... + wₙ·xₙ + b

The weights wᵢ say how much y moves when xᵢ moves by one unit, all else held constant. The intercept b is the prediction when every xᵢ = 0.

python
# Local: pip install scikit-learn numpy
import numpy as np
from sklearn.linear_model import LinearRegression

X = np.array([[1], [2], [3], [4], [5]])      # one feature, five samples
y = np.array([2.1, 4.0, 6.1, 7.9, 10.2])      # roughly y = 2x

model = LinearRegression()
model.fit(X, y)

print(model.coef_)                            # [2.03]   — slope ~2
print(model.intercept_)                       # 0.01      — intercept ~0
print(model.predict([[6]]))                   # [12.19]   — extrapolation
print(model.score(X, y))                      # 0.9994    — R²

fit(X, y) solves the normal equations under the hood — closed-form, no iteration. predict applies the learned weights. score returns R² (coefficient of determination) by default.


3. Multiple Features and Coefficient Interpretation

python
import numpy as np
from sklearn.linear_model import LinearRegression

# rows: samples; cols: [rooms, age_years, sqft]
X = np.array([[3, 10, 1500],
              [4, 25, 2200],
              [2, 40,  900],
              [5,  5, 2800],
              [3, 15, 1700]])
y = np.array([420_000, 530_000, 220_000, 740_000, 460_000])

model = LinearRegression().fit(X, y)
print(dict(zip(["rooms", "age", "sqft"], model.coef_)))
# {'rooms': 12_530.4, 'age': -2_870.1, 'sqft': 198.7}

Read coefficients carefully: "holding age and sqft fixed, one extra room is worth ~£12.5k". Two pitfalls:

  • Scale dependence — sqft has tiny coefficient but huge range; rooms has big coefficient but small range. You cannot rank features by raw coefficient magnitude unless you've standardised the inputs first (mean 0, std 1) — see feature.
  • Causation — a coefficient is a conditional association, not a causal effect. Adding a sunroom doesn't cause the price bump if your data merely correlates the two.

4. Train/Test Split — The Non-Negotiable

Evaluating a model on the data it learned from is like grading a student on the exam they wrote. You'll overestimate performance every time.

python
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

model = LinearRegression().fit(X_train, y_train)
print(model.score(X_test, y_test))            # honest R² on unseen data
+ setup added so this can run · defines X, y, LinearRegression
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X = _AutoMock('X')
y = _AutoMock('y')
def LinearRegression(*_a, **_kw):
    print('-> LinearRegression() called')
    return _AutoMock('LinearRegression()')
  • test_size=0.2 — hold out 20% for evaluation.
  • random_state=42 — reproducible split. Pick any integer; use the same one across the team.
  • For time-series, never random-split — use TimeSeriesSplit so the test set is always after training in time.

5. Regression Metrics — When to Use Which

python
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
import numpy as np

y_pred = model.predict(X_test)

mae  = mean_absolute_error(y_test, y_pred)
mse  = mean_squared_error(y_test, y_pred)
rmse = np.sqrt(mse)
r2   = r2_score(y_test, y_pred)
+ setup added so this can run · defines X_test, y_test, model
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_test = _AutoMock('X_test')
y_test = _AutoMock('y_test')
model = _AutoMock('model')
MetricWhat it measuresUse when
MAEAverage absolute error, same units as yYou want robust, intuitive "off by £X on average"
MSEAverage squared errorPenalises big misses harder; the optimisation target
RMSE√MSE, back in original unitsReporting; the de-facto industry metric
R²Fraction of variance explained, in [-∞, 1]Comparing to a "predict the mean" baseline

R² > 0 means you beat the mean baseline; R² = 1 is perfect; R² can go negative if your model is worse than predicting the mean. RMSE > MAE always — if they diverge wildly, you have large outliers.


6. The Textbook Assumptions (and Reality)

Statistics textbooks insist linear regression assumes:

1. Linearity — the true relationship is linear in the inputs.
2. Independence — observations don't influence each other.
3. Homoscedasticity — error variance is constant across x.
4. Normality of residuals — errors are Gaussian.
5. No multicollinearity — features aren't redundant.

In practice, no production team checks all five. We check residual plots for obvious structure (curvature → add polynomial features; funnel shape → log-transform y), eyeball collinearity via correlation matrices or VIF, and rely on cross-validated test-set RMSE to tell us whether the model is useful. Worry about the assumptions when inference matters (p-values, confidence intervals); for pure prediction, generalisation error is king.


7. Regularisation — Ridge, Lasso, ElasticNet

When you have many features (or few samples), plain linear regression overfits — it memorises noise. Regularisation adds a penalty on coefficient size, shrinking weights toward zero.

python
from sklearn.linear_model import Ridge, Lasso, ElasticNet
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline

# Regularised models REQUIRE scaled features — coefficient magnitude is now meaningful
ridge = Pipeline([
    ("scale", StandardScaler()),
    ("model", Ridge(alpha=1.0)),              # L2 — shrinks all coefs smoothly
])

lasso = Pipeline([
    ("scale", StandardScaler()),
    ("model", Lasso(alpha=0.1)),              # L1 — drives some coefs to ZERO (feature selection)
])

enet = Pipeline([
    ("scale", StandardScaler()),
    ("model", ElasticNet(alpha=0.1, l1_ratio=0.5)),   # mix of L1 and L2
])
ModelPenaltyEffectWhen
Ridgeα · Σ wᵢ²Shrinks coefs smoothly, keeps all featuresMany correlated features; you want stability
Lassoα · Σ \|wᵢ\|Drives weak coefs to exactly 0You want automatic feature selection
ElasticNetmix of bothBest of bothCorrelated features + selection wanted

alpha is the regularisation strength — bigger means more shrinkage. Tune it via cross-validation (Section 13).

Why scaling is mandatory: the penalty Σ wᵢ² punishes large weights. A feature on scale 0–1,000,000 will need a tiny weight to make any contribution and won't get penalised; a feature on scale 0–1 will need a huge weight and get crushed. Without scaling, regularisation is biased toward whichever features happen to have small numeric ranges.


8. Logistic Regression — Same Idea, Different Output

For binary classification (e.g. spam vs ham), compute the same linear combination and squash through the sigmoid:

python
z = w₁·x₁ + ... + wₙ·xₙ + b
p = 1 / (1 + e⁻ᶻ)              # in (0, 1) — interpret as P(class = 1)

Predict class 1 if p ≥ 0.5, else class 0. The 0.5 boundary is convention, not law (more on threshold tuning below).

python
from sklearn.linear_model import LogisticRegression
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split

X, y = make_classification(n_samples=1000, n_features=10, random_state=42)
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

clf = LogisticRegression(max_iter=1000).fit(X_train, y_train)

print(clf.predict(X_test[:5]))               # [0 1 1 0 1]
print(clf.predict_proba(X_test[:5]))         # probabilities per class
print(clf.score(X_test, y_test))             # accuracy

Coefficients of logistic regression are log-odds: a one-unit increase in xᵢ multiplies the odds of class 1 by exp(wᵢ). Useful for inference; almost nobody bothers in pure prediction work.


9. Threshold Tuning

The default 0.5 threshold is fine when classes are balanced and false positives/negatives cost the same. They almost never do.

python
import numpy as np
probs = clf.predict_proba(X_test)[:, 1]      # P(class = 1)

# Aggressive — flag anything with > 30% probability (high recall, more false positives)
y_pred_aggressive = (probs > 0.30).astype(int)

# Conservative — only flag if > 70% confident (high precision, more false negatives)
y_pred_conservative = (probs > 0.70).astype(int)
+ setup added so this can run · defines X_test, clf
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_test = _AutoMock('X_test')
clf = _AutoMock('clf')

Cancer screening tolerates false positives (re-test); fraud blocking tolerates false negatives (manual review queue) — pick the threshold by cost, not by 0.5 reflex.


10. Classification Metrics — Why Accuracy Lies

python
from sklearn.metrics import (
    accuracy_score, precision_score, recall_score, f1_score,
    roc_auc_score, confusion_matrix, classification_report
)

y_pred = clf.predict(X_test)
print(confusion_matrix(y_test, y_pred))
# [[TN FP]
#  [FN TP]]

print(classification_report(y_test, y_pred))
#               precision    recall  f1-score   support
#            0       0.91      0.89      0.90       100
#            1       0.89      0.91      0.90       100
#     accuracy                           0.90       200
+ setup added so this can run · defines X_test, clf, y_test
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_test = _AutoMock('X_test')
clf = _AutoMock('clf')
y_test = _AutoMock('y_test')
MetricFormulaUse when
Accuracy(TP+TN) / totalBalanced classes, equal costs
PrecisionTP / (TP+FP)False positives are expensive (spam filter — don't flag real mail)
RecallTP / (TP+FN)False negatives are expensive (cancer — don't miss a case)
F1harmonic mean of P and RBalance both; default for imbalanced
ROC-AUCarea under ROC curveRanking quality across all thresholds

Accuracy lies on imbalanced data. If 99% of transactions are legitimate, a model that predicts "legitimate" for everything scores 99% accuracy and catches zero fraud. Always check the confusion matrix; for imbalanced problems, optimise F1 or AUC (or precision/recall at a chosen threshold tied to business cost).


11. ROC-AUC in One Snippet

python
from sklearn.metrics import roc_auc_score, roc_curve

probs = clf.predict_proba(X_test)[:, 1]
auc = roc_auc_score(y_test, probs)
print(auc)                                    # 0.95 — closer to 1 is better; 0.5 is random
fpr, tpr, thresholds = roc_curve(y_test, probs)
+ setup added so this can run · defines y_test, X_test, clf
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

y_test = _AutoMock('y_test')
X_test = _AutoMock('X_test')
clf = _AutoMock('clf')

AUC = 0.5 is coin-flip; 1.0 is perfect. It measures the probability that a random positive is ranked higher than a random negative — threshold-free, so it's the right summary when you haven't picked a threshold yet.


12. Cross-Validation — One Split Isn't Enough

A single train/test split is one sample of model performance. The variance across splits can be large — your "85% accuracy" might be 78–92% depending on the seed. k-fold cross-validation runs k train/test splits across the dataset and averages the scores.

python
from sklearn.model_selection import cross_val_score

scores = cross_val_score(LogisticRegression(max_iter=1000), X, y, cv=5, scoring="f1")
print(scores)                                 # [0.89 0.91 0.87 0.90 0.88]
print(scores.mean(), scores.std())            # 0.89 ± 0.014
+ setup added so this can run · defines X, y, LogisticRegression
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X = _AutoMock('X')
y = _AutoMock('y')
def LogisticRegression(*_a, **_kw):
    print('-> LogisticRegression() called')
    return _AutoMock('LogisticRegression()')

cv=5 is the default sweet spot. Use StratifiedKFold (sklearn's default for classifiers) so each fold preserves the class ratio. For time-series, TimeSeriesSplit. Report mean ± std — a single number hides the variance.


13. Hyperparameter Tuning with GridSearchCV

python
from sklearn.model_selection import GridSearchCV
from sklearn.linear_model import Ridge
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline

pipe = Pipeline([("scale", StandardScaler()), ("model", Ridge())])

param_grid = {"model__alpha": [0.001, 0.01, 0.1, 1, 10, 100]}

search = GridSearchCV(pipe, param_grid, cv=5, scoring="neg_root_mean_squared_error")
search.fit(X_train, y_train)

print(search.best_params_)                    # {'model__alpha': 1.0}
print(-search.best_score_)                    # 12_340.5 (RMSE)
print(search.score(X_test, y_test))           # held-out test score
+ setup added so this can run · defines X_train, y_train, X_test, y_test
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_train = _AutoMock('X_train')
y_train = _AutoMock('y_train')
X_test = _AutoMock('X_test')
y_test = _AutoMock('y_test')

Key points:

  • model__alpha — the double underscore selects the alpha parameter on the step named model inside the pipeline.
  • Scoring is negated (neg_root_mean_squared_error) because GridSearchCV always maximises — for losses you want minimised, flip the sign.
  • Always fit the pipeline inside the CV loop — scaling-before-split leaks test statistics into training.

For large grids, RandomizedSearchCV samples instead of exhaustively trying every combination.


14. Multi-Class Logistic

Three or more classes? Two strategies:

python
LogisticRegression(multi_class="ovr")          # one-vs-rest — train K binary classifiers
LogisticRegression(multi_class="multinomial")  # true softmax — joint probability over K classes
+ setup added so this can run · defines LogisticRegression
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def LogisticRegression(*_a, **_kw):
    print('-> LogisticRegression() called')
    return _AutoMock('LogisticRegression()')

multinomial (softmax) is usually more accurate and well-calibrated; ovr is faster to train and easier to explain. Sklearn picks multinomial automatically with the lbfgs solver for most datasets.


Common Mistakes

1. Not scaling features for regularised regression. Ridge/Lasso/ElasticNet penalties depend on coefficient magnitude. Unscaled features → biased penalty → wrong coefs. Always pipeline StandardScaler before the model. Plain LinearRegression doesn't care.

2. Accuracy on imbalanced data.

python
# 99% class 0, 1% class 1 — predict-everything-zero
y_pred = np.zeros_like(y_test)
print(accuracy_score(y_test, y_pred))         # 0.99 — meaningless
print(recall_score(y_test, y_pred))           #   0   — caught zero positives
+ setup added so this can run · defines y_test, np, accuracy_score, recall_score
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

y_test = _AutoMock('y_test')
np = _AutoMock('np')
def accuracy_score(*_a, **_kw):
    print('-> accuracy_score() called')
    return _AutoMock('accuracy_score()')
def recall_score(*_a, **_kw):
    print('-> recall_score() called')
    return _AutoMock('recall_score()')

Use F1, AUC, or recall@precision-floor instead. Set class_weight="balanced" to make the loss attend to the minority class.

3. Predicting on un-transformed data. If you scaled X_train with StandardScaler, you must apply the same fitted scaler to X_test (and any new prediction input). Easiest fix: wrap everything in a Pipeline — fit-transform on train, transform-only on test happens automatically.

4. Data leakage via pre-split transformation.

python
# WRONG — scaler sees test statistics
scaler = StandardScaler().fit(X)              # fit on EVERYTHING
X_scaled = scaler.transform(X)
X_train, X_test, y_train, y_test = train_test_split(X_scaled, y, ...)
+ setup added so this can run · defines X, train_test_split, y, StandardScaler
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X = _AutoMock('X')
def train_test_split(*_a, **_kw):
    print('-> train_test_split() called')
    return _AutoMock('train_test_split()')
y = _AutoMock('y')
def StandardScaler(*_a, **_kw):
    print('-> StandardScaler() called')
    return _AutoMock('StandardScaler()')

The scaler has now memorised the test set's mean and std — your reported test score is optimistic. Always split first, then fit transformers on X_train only. Pipelines + cross-validation do this correctly out of the box.

5. Coefficient hunting. Reading raw model.coef_ without (a) standardising features or (b) statistical significance tests is a beginner trap. For ranking feature importance honestly, use permutation importance (covered in trees).

6. max_iter warnings. LogisticRegression defaults to 100 iterations; large/unscaled problems hit the cap before converging — sklearn warns. Either scale the data or bump to max_iter=1000.


🎯 Your Turn — A Full Ridge Pipeline

Build a complete regression workflow on the inline boston_lite data: split, scale, fit Ridge with α tuned via 5-fold CV, predict on the held-out set, report RMSE and R².

Requirements:

1. Use train_test_split with test_size=0.2, random_state=42.
2. Wrap StandardScaler and Ridge in a Pipeline.
3. Use GridSearchCV to tune alpha over [0.01, 0.1, 1, 10, 100] with 5-fold CV.
4. Report best_params_, the test-set RMSE, and the test-set R².

Skeleton:

python
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.metrics import mean_squared_error, r2_score

# Inline mini dataset: [rooms, age, lstat] -> median house value (in $1000s)
X = np.array([
    [6.575, 65.2, 4.98], [6.421, 78.9, 9.14], [7.185, 61.1, 4.03],
    [6.998, 45.8, 2.94], [7.147, 54.2, 5.33], [6.430, 58.7, 5.21],
    [6.012, 66.6, 12.43], [6.172, 96.1, 19.15], [5.631, 100.0, 29.93],
    [6.004, 85.9, 17.10], [6.377, 70.3, 20.45], [6.009, 94.3, 13.27],
    [5.889, 39.0, 15.71], [7.249, 21.9, 4.81], [6.625, 17.5, 3.81],
    [5.927, 42.6, 18.66], [6.137, 32.1, 13.44], [5.951, 38.5, 12.13],
    [6.456, 71.9, 11.10], [6.815, 40.5, 5.25],
])
y = np.array([24.0, 21.6, 34.7, 33.4, 36.2, 28.7, 22.9, 17.5, 16.5, 18.9,
              15.0, 18.9, 21.7, 35.4, 28.7, 19.6, 23.1, 21.7, 22.2, 30.8])

# TODO 1: split into train/test (test_size=0.2, random_state=42)
# TODO 2: build Pipeline([("scale", StandardScaler()), ("model", Ridge())])
# TODO 3: GridSearchCV over model__alpha with cv=5, scoring="neg_root_mean_squared_error"
# TODO 4: fit on training data, predict on test
# TODO 5: print best_params_, RMSE, R²
Hint 1 — Pipeline parameter names Inside GridSearchCV, parameters belonging to a pipeline step are named stepname__paramname — two underscores. So Ridge's alpha inside a step called "model" becomes "model__alpha". Get this wrong and you'll see a baffling ValueError: Invalid parameter.
Hint 2 — RMSE from MSE mean_squared_error returns MSE by default. Either pass squared=False (sklearn ≥ 0.22) to get RMSE directly, or take np.sqrt(mean_squared_error(...)). Also remember GridSearchCV's best_score_ is negated — flip its sign before reporting.
Show full solution
python
import numpy as np
from sklearn.linear_model import Ridge
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import Pipeline
from sklearn.model_selection import train_test_split, GridSearchCV
from sklearn.metrics import mean_squared_error, r2_score

X = np.array([
    [6.575, 65.2, 4.98], [6.421, 78.9, 9.14], [7.185, 61.1, 4.03],
    [6.998, 45.8, 2.94], [7.147, 54.2, 5.33], [6.430, 58.7, 5.21],
    [6.012, 66.6, 12.43], [6.172, 96.1, 19.15], [5.631, 100.0, 29.93],
    [6.004, 85.9, 17.10], [6.377, 70.3, 20.45], [6.009, 94.3, 13.27],
    [5.889, 39.0, 15.71], [7.249, 21.9, 4.81], [6.625, 17.5, 3.81],
    [5.927, 42.6, 18.66], [6.137, 32.1, 13.44], [5.951, 38.5, 12.13],
    [6.456, 71.9, 11.10], [6.815, 40.5, 5.25],
])
y = np.array([24.0, 21.6, 34.7, 33.4, 36.2, 28.7, 22.9, 17.5, 16.5, 18.9,
              15.0, 18.9, 21.7, 35.4, 28.7, 19.6, 23.1, 21.7, 22.2, 30.8])

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

pipe = Pipeline([
    ("scale", StandardScaler()),
    ("model", Ridge()),
])

param_grid = {"model__alpha": [0.01, 0.1, 1.0, 10.0, 100.0]}

search = GridSearchCV(
    pipe, param_grid, cv=5,
    scoring="neg_root_mean_squared_error",
)
search.fit(X_train, y_train)

y_pred = search.predict(X_test)
rmse = np.sqrt(mean_squared_error(y_test, y_pred))
r2   = r2_score(y_test, y_pred)

print(f"Best alpha: {search.best_params_['model__alpha']}")
print(f"CV RMSE:    {-search.best_score_:.3f}")
print(f"Test RMSE:  {rmse:.3f}")
print(f"Test R²:    {r2:.3f}")

# Expected output (approximate — small dataset, results vary slightly):
# Best alpha: 1.0
# CV RMSE:    4.1
# Test RMSE:  3.8
# Test R²:    0.72

Why this solution is right by construction:

  • Split first — the scaler never sees X_test, no leakage.
  • Pipeline — the same scaler that fit on X_train is automatically applied at predict time on X_test. No risk of forgetting to transform.
  • Cross-validated tuning — alpha=1.0 is picked because it gave the best mean RMSE across the 5 folds of the training set, not because of one lucky split.
  • Final evaluation on held-out test — the only honest performance number is the one on data that touched neither the scaler's .fit() nor the alpha search.

Drop in Lasso instead of Ridge and you get automatic feature selection — coefficients of weak features become exactly zero. Same pipeline, one-word change.


What You Learned

  • Regression predicts a number; classification predicts a category. Linear and logistic regression are the canonical starting models for each.
  • Linear regression fits y = w·x + b by minimising squared errors; coefficients are interpretable only after standardising features.
  • train_test_split with a fixed random_state gives you an honest, reproducible held-out evaluation.
  • Regression metrics: RMSE for industry reporting, MAE for robustness to outliers, R² to compare against the mean baseline.
  • Regularisation — Ridge (L2) shrinks smoothly; Lasso (L1) zeroes out weak features; ElasticNet mixes both. All three require scaled inputs.
  • Logistic regression = linear combination → sigmoid → probability → threshold. Default 0.5 is convention, not law.
  • Classification metrics: accuracy lies on imbalanced data — use F1, AUC, or precision/recall tied to business cost.
  • k-fold cross-validation beats a single split for estimating performance and tuning hyperparameters.
  • GridSearchCV + Pipeline is the safe, leakage-free pattern for hyperparameter tuning.
  • Multi-class logistic uses one-vs-rest or softmax (multinomial); sklearn picks sensibly by default.

Next: Decision Trees & Random Forests — non-linear models, ensembles, and the trees that beat linear baselines on most tabular data.