Linear Regression: Predicting Numbers
1 · The lesson
readYou're predicting a number. Sale price. Number of clicks. Tomorrow's temperature. The target is continuous — it could be 47, 47.3, or 47.291.
Reach for linear regression first. It's the oldest tool in the box, it trains in milliseconds, the coefficients are interpretable, and on a surprising number of real problems it beats fancier models. If linear regression fails, you've also learned something — the relationship probably isn't linear, and you've earned the right to try something heavier.
This lesson is about that "first move".
Run these right here — scikit-learn, pandas and matplotlib all work in the browser. The first scikit-learn import takes a few seconds while it downloads; after that it's instant. Expected output is also shown in comments below each block.
1. When Linear Regression Is the Right Tool
Three boxes to tick:
- Target is continuous — a real number, not a category.
- You expect a roughly linear relationship — bigger house, higher price; more ad spend, more sales. Not "U-shaped" or "explodes after a threshold".
- You want to interpret the model — "for every extra bedroom, price goes up by £30k". Linear coefficients tell you exactly that.
If the target is a category (spam/not-spam), you want classification, not regression. If the relationship is wildly non-linear, you'll want trees.
2. The Mental Model
You have a cloud of points. Linear regression draws the single straight line that gets closest to all of them. "Closest" means minimum sum of squared distances — which is fancy talk for "the line that passes through the middle of the cloud".
price ▲ ● │ ● ● │ ●● ● │ ● ● ● │ ● ●●● │ ●●●● ╲ ← the line │●●● (best fit) └────────────────────────► square footage
That's the whole intuition. One feature → one line. Two features → a tilted plane in 3D. Many features → a "hyperplane" you can't draw but the maths handles fine.
3. The Equation Behind It
One feature:
y = m·x + b
setup added so this can run · defines m·x, b
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) m·x = _AutoMock('m·x') b = _AutoMock('b')
Same y = mx + b you saw in school. m is the slope (how much y changes when x goes up by 1), b is the intercept (y when x is 0). Linear regression's only job is to learn the best m and b from data.
Multiple features:
y = w₁·x₁ + w₂·x₂ + w₃·x₃ + ... + b
Each feature gets its own weight (w_i, also called a coefficient). The model learns one weight per feature, plus one intercept. Still a line, just in more dimensions.
You won't write this maths yourself. scikit-learn fits the weights for you in a single .fit() call.
4. Code: House Prices in 12 Lines
import numpy as np import pandas as pd from sklearn.linear_model import LinearRegression from sklearn.model_selection import train_test_split df = pd.DataFrame({ "sqft": [800, 1200, 1500, 1800, 2200, 2600, 3000, 3500], "bedrooms": [1, 2, 2, 3, 3, 4, 4, 5], "price": [180, 250, 300, 370, 430, 520, 590, 690], # £k }) X = df[["sqft", "bedrooms"]] y = df["price"] X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42) model = LinearRegression().fit(X_train, y_train) print("Coefficients:", model.coef_) print("Intercept: ", model.intercept_) print("R²: ", model.score(X_test, y_test)) # → Coefficients: [0.171 18.62] # → Intercept: 34.5 # → R²: 0.998
The model learned: price ≈ 0.171·sqft + 18.62·bedrooms + 34.5 (in £k).
Translate that:
- Every extra square foot adds about £171 to the price.
- Every extra bedroom adds about £18,620 — holding sqft constant.
- A 0-sqft, 0-bedroom house would cost £34,500. (Nonsense, of course — that's the intercept just doing its arithmetic job.)
R² of 0.998 means the model explains 99.8% of the variance in price on the test set. With 8 rows of clean synthetic data, that's expected. Real data is messier.
5. Interpreting Coefficients — the Big Caveat
The "every extra bedroom adds £18,620" reading only works when features are on comparable scales. Sqft ranges from 800 to 3500. Bedrooms ranges from 1 to 5. The coefficients look incomparable because the units are incomparable.
To compare relative importance, scale the features first:
from sklearn.preprocessing import StandardScaler scaler = StandardScaler() X_scaled = scaler.fit_transform(X_train) model = LinearRegression().fit(X_scaled, y_train) print(dict(zip(X.columns, model.coef_))) # → {'sqft': 142.7, 'bedrooms': 23.1}
setup added so this can run · defines X_train, y_train, LinearRegression, X
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_train = _AutoMock('X_train') y_train = _AutoMock('y_train') def LinearRegression(*_a, **_kw): print('-> LinearRegression() called') return _AutoMock('LinearRegression()') X = _AutoMock('X')
Now the coefficients live in the same "1 standard deviation" units, and you can honestly say sqft has roughly 6× the impact of bedroom-count on price.
The fix-this-properly approach is to do scaling inside a Pipeline so it can't accidentally leak between train and test. That's lesson 6.
6. Metrics for Regression
.score() gives you R². That's a fine summary number, but you usually want at least one error metric in the units of your target.
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score import numpy as np y_pred = model.predict(X_test) mae = mean_absolute_error(y_test, y_pred) rmse = np.sqrt(mean_squared_error(y_test, y_pred)) r2 = r2_score(y_test, y_pred) print(f"MAE = £{mae:,.0f}k") print(f"RMSE = £{rmse:,.0f}k") print(f"R² = {r2:.3f}") # → MAE = £4k # → RMSE = £5k # → R² = 0.998
setup added so this can run · defines X_test, y_test, model
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_test = _AutoMock('X_test') y_test = _AutoMock('y_test') model = _AutoMock('model')
| Metric | Reads as | When to prefer it |
|---|---|---|
| MAE | "Average error of £X in the target unit." | Easy to explain to a stakeholder. |
| RMSE | "Like MAE, but punishes big misses more." | When occasional huge errors are unacceptable. |
| R² | "% of variance the model explains." | Sanity-check the model is doing anything. |
A common interview question: "why is RMSE always ≥ MAE?" — because squaring before averaging weights big errors more, then sqrt-ing brings you back to original units but with the bias baked in.
7. When Linear Regression Fails
The model's only knob is a slope per feature. If the real relationship is curved, it can't bend.
y ▲ ●●● │ ●● ●● │ ●● ●● ← real pattern (curved) │ ● ● │ ● ● ── the best straight line LR can do │● ● └─────────────────► x
Two escape hatches:
1. Polynomial features. Manually build x², x³ columns; linear regression on those is curve-fitting.
from sklearn.preprocessing import PolynomialFeatures from sklearn.pipeline import make_pipeline model = make_pipeline(PolynomialFeatures(degree=3), LinearRegression())
setup added so this can run · defines LinearRegression
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def LinearRegression(*_a, **_kw): print('-> LinearRegression() called') return _AutoMock('LinearRegression()')
2. Switch to trees. Decision trees and random forests handle non-linearity natively, no feature-engineering needed. They lose the interpretability though.
How do you know you have this problem? Plot the residuals.
8. The Residual Plot — Your Diagnostic
A residual is actual - predicted for each point. If linear regression fits well, residuals should be a structureless cloud around zero. If there's a pattern in the residuals, you've left signal on the table.
import matplotlib.pyplot as plt y_pred = model.predict(X_test) residuals = y_test - y_pred plt.scatter(y_pred, residuals) plt.axhline(0, color="red", linestyle="--") plt.xlabel("Predicted price"); plt.ylabel("Residual (actual - predicted)") plt.title("Residual plot — looking for randomness") plt.show()
setup added so this can run · defines X_test, y_test, model
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_test = _AutoMock('X_test') y_test = _AutoMock('y_test') model = _AutoMock('model')
What you want to see: random scatter around 0. What screams "non-linear pattern": a clear curve or fan-shape in the residuals. That's your signal to try polynomial features or a tree-based model.
Common Mistakes
- Using accuracy on a regression problem. Accuracy is for classification. You'll get either a
TypeErroror a meaningless 0.0. Use MAE / RMSE / R². - Interpreting unscaled coefficients as importance. A coefficient of 0.171 looks tiny next to 18.62 — until you realise one is per square foot and the other per bedroom. Scale first.
- Forgetting the linearity assumption. Linear regression on a clearly curved relationship will run, return a number, and quietly be wrong. Always plot residuals.
- Including the target as a feature. It happens. You drop columns by mistake and
priceends up inX. The model gets R² = 1.0 and you ship a useless model.
🎯 Your Turn — Salary vs Years of Experience
Inline data:
years = [1, 2, 3, 4, 5, 6, 7, 8, 9, 10] salary = [35, 42, 47, 53, 60, 64, 71, 77, 82, 89] # £k
Fit a linear regression. Then:
1. Predict the salary for 5 years of experience.
2. Report the RMSE on the full dataset (we're not splitting, only 10 rows).
3. Return (predicted_salary_for_5_years, rmse) rounded to 2 decimals.
import numpy as np from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error def fit_and_predict(): years = np.array([1, 2, 3, 4, 5, 6, 7, 8, 9, 10]).reshape(-1, 1) salary = np.array([35, 42, 47, 53, 60, 64, 71, 77, 82, 89]) # TODO 1: instantiate and fit LinearRegression # TODO 2: predict for years = 5 # TODO 3: compute RMSE on the full dataset # TODO 4: return (round(prediction, 2), round(rmse, 2)) pass print(fit_and_predict()) # → expected something near (59.95, 0.9x)
Hint 1 — predict expects 2D input
model.predict([[5]]) — note the double brackets. predict wants a 2D array even for one prediction. The result is also a 1D array, so use [0] to extract the scalar.
Hint 2 — RMSE in one line
rmse = np.sqrt(mean_squared_error(salary, model.predict(years))). Compare actuals to predictions on the same data the model was trained on (fine here because we're only checking fit quality, not generalisation).
Show full solution
import numpy as np from sklearn.linear_model import LinearRegression from sklearn.metrics import mean_squared_error def fit_and_predict(): years = np.array([1, 2, 3, 4, 5, 6, 7, 8, 9, 10]).reshape(-1, 1) salary = np.array([35, 42, 47, 53, 60, 64, 71, 77, 82, 89]) model = LinearRegression().fit(years, salary) prediction = model.predict([[5]])[0] rmse = np.sqrt(mean_squared_error(salary, model.predict(years))) return (round(prediction, 2), round(rmse, 2)) print(fit_and_predict()) # → (59.95, 0.91)
The model learned a slope of about £5.97k per year and an intercept of about £30k. The fit is tight — RMSE under £1k on salaries in the £35-£89k range. Linear regression earned its keep here.
What You Learned
- Linear regression predicts a continuous target by fitting one straight line (or hyperplane).
- The equation
y = w₁x₁ + ... + b— the model just learns the weights. LinearRegression().fit(X, y), then.coef_,.intercept_,.predict(),.score().- Interpret coefficients only after scaling features to comparable units.
- Three metrics: MAE (avg error), RMSE (penalises big misses), R² (variance explained).
- When residuals show a pattern, the relationship is non-linear — reach for polynomial features or trees.
Next: Classification — when the answer is a category, not a number.