PythonMastery
beginner 14 min read · lesson 2 of 6 in Machine Learning Fundamentals

Your Data, Your Model

1 · The lesson

read

Every supervised ML problem boils down to one contract: given X, predict y.

X is what the model sees. y is what you want it to predict. Get this shape right and 80% of scikit-learn just works. Get it wrong and you'll spend an afternoon wondering why a "shape mismatch" error keeps shouting at you.

This lesson is the bridge between "I have a CSV" and "I have a model that predicts something". The minimum viable ML script lives in this lesson.

Run these right here — scikit-learn, pandas and matplotlib all work in the browser. The first scikit-learn import takes a few seconds while it downloads; after that it's instant. Expected output is also shown in comments below each block.


1. Features and Target

The convention every ML library on Earth follows:

python
X = features  (the inputs, plural)
y = target    (the thing you predict, singular)

Imagine predicting house prices:

sqftbedroomsageprice
1200212320,000
180035480,000
950130210,000
  • X is the first three columns: sqft, bedrooms, age. Shape (3, 3) — three rows, three features.
  • y is the last column: price. Shape (3,) — one value per row.

In NumPy / pandas terms:

python
import pandas as pd

df = pd.DataFrame({
    "sqft":     [1200, 1800, 950],
    "bedrooms": [2,    3,    1],
    "age":      [12,   5,    30],
    "price":    [320_000, 480_000, 210_000],
})

X = df[["sqft", "bedrooms", "age"]]
y = df["price"]

print(X.shape, y.shape)
# → (3, 3) (3,)

The capital-X-lowercase-y convention is a hold-over from maths notation — X is a matrix, y is a vector. Every example you read uses it. Use it too.


2. The Train/Test Split

A model that "predicts" what it has already seen is useless — that's just a lookup table. The whole point is generalisation: doing well on data it has never seen.

So before training, you carve off a chunk of data and hide it. Train on the rest. Score on the hidden chunk. That score is your honest estimate of real-world performance.

python
from sklearn.model_selection import train_test_split

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)
+ setup added so this can run · defines X, y
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X = _AutoMock('X')
y = _AutoMock('y')
  • test_size=0.2 — keep 20% for testing.
  • random_state=42 — fix the shuffle so the split is reproducible. Use any integer; pick one and stick with it across the project.

After this line you have four arrays:

python
X_train, y_train  →  model sees these during .fit()
X_test,  y_test   →  model never sees these until evaluation

If you train on test data, your score is meaningless. This is the cardinal sin of ML — don't grade your own homework.


3. The sklearn API in Three Methods

Every scikit-learn estimator — and there are hundreds — has the same three methods:

python
model.fit(X_train, y_train)      # learn from data
model.predict(X_test)            # make predictions on new data
model.score(X_test, y_test)      # one-number performance metric
+ setup added so this can run · defines X_train, y_train, X_test, y_test, model
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_train = _AutoMock('X_train')
y_train = _AutoMock('y_train')
X_test = _AutoMock('X_test')
y_test = _AutoMock('y_test')
model = _AutoMock('model')

That's the entire API contract. Once you've written this for one model, you've written it for all of them. Swap LinearRegression() for RandomForestRegressor() and the surrounding code is identical.

.score() returns a sensible default metric:


  • For classifiers, it's accuracy.

  • For regressors, it's R² (variance explained, 1.0 is perfect, 0.0 means "no better than predicting the mean").

You'll often want better metrics than .score() — we cover those in lesson 4 — but it's a fine first pass.


4. The Minimum Viable ML Script

Ten lines. Memorise the shape.

python
from sklearn.datasets import load_iris
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression

X, y = load_iris(return_X_y=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, random_state=42
)

model = LogisticRegression(max_iter=1000)
model.fit(X_train, y_train)
print(f"Test accuracy: {model.score(X_test, y_test):.2f}")
# → Test accuracy: 1.00

Look at the structure:

1. Load the data.
2. Split it.
3. Instantiate the model.
4. Fit on the training data.
5. Score on the test data.

This shape works for 95% of classical ML problems. You'll change the model, the data, the metric — the skeleton stays the same.


5. Underfit, Overfit, and the Goldilocks Zone

When a model performs badly, it's failing in one of two ways.

Underfitting — too simple

The model is so weak it can't even capture the training data. Training error is high, test error is high. Imagine drawing a straight line through a clearly curved cloud of points.

Overfitting — too complex

The model memorises the training data, including the noise. Training error is tiny, test error is huge. Imagine drawing a wiggly line that passes through every single training point — perfectly fitted to the past, useless for the future.

The classic visual: fitting a noisy sine wave with polynomials of different degrees.

python
import numpy as np
import matplotlib.pyplot as plt
from sklearn.linear_model import LinearRegression
from sklearn.preprocessing import PolynomialFeatures
from sklearn.pipeline import make_pipeline

rng = np.random.default_rng(42)
X = np.linspace(0, 2 * np.pi, 30).reshape(-1, 1)
y = np.sin(X).ravel() + rng.normal(0, 0.2, size=30)

for degree, label in [(1, "underfit"), (4, "just right"), (15, "overfit")]:
    model = make_pipeline(PolynomialFeatures(degree), LinearRegression())
    model.fit(X, y)
    plt.scatter(X, y, alpha=0.5)
    grid = np.linspace(0, 2 * np.pi, 200).reshape(-1, 1)
    plt.plot(grid, model.predict(grid), label=f"deg {degree} — {label}")

plt.legend(); plt.title("Underfit / Just right / Overfit"); plt.show()

You'd see:


  • Degree 1 — a straight line through a curve. Underfit.

  • Degree 4 — follows the sine smoothly. Just right.

  • Degree 15 — wiggles through every point. Overfit.

The goldilocks zone is whichever complexity gives you the lowest test error. That's why we have a test set in the first place — to find it.


6. Validation Set vs Test Set

For real projects, two splits aren't enough. You want three roles:

  • Train — fit the model.
  • Validation — tune hyperparameters (the model's knobs).
  • Test — final, untouched score. Only look at this once.

If you tune knobs using the test set, the test set becomes part of your training process and stops being honest. You've leaked.

The cheap version is two train_test_split calls:

python
X_temp, X_test, y_temp, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
X_train, X_val, y_train, y_val = train_test_split(X_temp, y_temp, test_size=0.25, random_state=42)
# Now: 60% train, 20% validation, 20% test
+ setup added so this can run · defines train_test_split, X, y
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def train_test_split(*_a, **_kw):
    print('-> train_test_split() called')
    return _AutoMock('train_test_split()')
X = _AutoMock('X')
y = _AutoMock('y')

Cross-validation — when data is scarce

If you only have 200 rows, carving 40% off for validation+test stings. K-fold cross-validation lets every row take a turn as validation.

python
from sklearn.model_selection import cross_val_score

scores = cross_val_score(model, X_train, y_train, cv=5)
print(f"CV mean: {scores.mean():.3f}, std: {scores.std():.3f}")
# → CV mean: 0.967, std: 0.021
+ setup added so this can run · defines model, X_train, y_train
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

model = _AutoMock('model')
X_train = _AutoMock('X_train')
y_train = _AutoMock('y_train')
  • cv=5 — split training data into 5 folds.
  • Train on 4, score on 1. Repeat 5 times so every fold gets a turn.
  • Return 5 scores. The mean is your honest estimate; the std tells you how stable it is.

Cross-validation is the default move for any project with under ~10,000 rows. We use it heavily in pipelines.


Common Mistakes

  • Training on the test set. Sometimes accidental: you call train_test_split after fitting a scaler on all the data. The test set was already "seen" through the scaler's statistics. Lesson 6 fixes this with Pipeline.
  • Predicting on un-preprocessed data. Whatever you did to X_train (scaling, encoding, imputing) must happen to X_test with the train-fitted transformer. Forgetting this is the most common production bug in ML.
  • Reporting the training score. "My model is 99% accurate!" — on what? If it's training accuracy, you've told us nothing about generalisation.
  • Not setting random_state. Your colleague reruns your notebook, gets a different split, gets a different number, files a bug. Pin the seed.

🎯 Your Turn — Train Your First Classifier

Fill in the gaps to:

1. Generate a synthetic 2-class dataset.
2. Split it 80/20.
3. Fit any sklearn classifier of your choice.
4. Return (train_accuracy, test_accuracy) as a tuple, both rounded to 3 decimals.

python
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression  # or any other

def train_and_score():
    X, y = make_classification(n_samples=500, n_features=10,
                               n_informative=5, random_state=42)

    # TODO 1: split X, y into train/test (80/20), random_state=42
    # TODO 2: instantiate a classifier
    # TODO 3: fit on training data
    # TODO 4: compute train_acc and test_acc using model.score(...)
    # TODO 5: return (round(train_acc, 3), round(test_acc, 3))

    pass

print(train_and_score())
# → expected something like (0.918, 0.870)
Hint 1 — the split train_test_split(X, y, test_size=0.2, random_state=42) returns four arrays in the order X_train, X_test, y_train, y_test.
Hint 2 — two .score() calls model.score(X_train, y_train) gives training accuracy. model.score(X_test, y_test) gives test accuracy. Both return a float between 0 and 1.
Show full solution
python
from sklearn.datasets import make_classification
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LogisticRegression

def train_and_score():
    X, y = make_classification(n_samples=500, n_features=10,
                               n_informative=5, random_state=42)

    X_train, X_test, y_train, y_test = train_test_split(
        X, y, test_size=0.2, random_state=42
    )

    model = LogisticRegression(max_iter=1000)
    model.fit(X_train, y_train)

    train_acc = model.score(X_train, y_train)
    test_acc  = model.score(X_test,  y_test)

    return (round(train_acc, 3), round(test_acc, 3))

print(train_and_score())
# → (0.918, 0.870)

Training accuracy slightly higher than test is normal and healthy. If they were identical, your model might be too simple. If training was 0.99 and test was 0.65, you'd be overfitting. The small gap here is the sweet spot.


What You Learned

  • The contract: X (features) and y (target). Capital X, lowercase y. Always.
  • train_test_split carves off a hold-out chunk you never touch during training.
  • Every sklearn estimator has .fit(), .predict(), .score(). Same shape, different models.
  • The minimum viable ML script: load → split → fit → score. Five steps.
  • Underfit = too simple. Overfit = memorised noise. Goldilocks = best test score.
  • Use a validation set (or cross-validation) for tuning, keep the test set untouched until the end.

Next: Linear Regression — predicting numbers with the simplest model in the toolkit.