Your First ML Pipeline
1 · The lesson
readReal ML code is fragile. You scale the training data, fit a model, then forget to scale the test data. Or you scale all the data before splitting, leaking the test set's statistics into training. Or you save just the model and discover at deployment time that you never serialised the encoder.
Every one of these is a bug you can't see — the model still runs, the numbers still look plausible, the predictions are quietly wrong.
scikit-learn's Pipeline is the fix. It chains preprocessing steps and the final model into a single estimator that fits, predicts, and serialises as one unit. Build this habit on day one and you'll save yourself a year of debugging.
Run these right here — scikit-learn, pandas and joblib all work in the browser. The first scikit-learn import takes a few seconds while it downloads. The one block that reads
data.csvis a template for your own data, so it needs a CSV of your own. Expected output is also shown in comments below each block.
1. The Problem Pipelines Solve
Here's the fragile version. Spot the bugs.
# DON'T do this from sklearn.preprocessing import StandardScaler from sklearn.linear_model import LogisticRegression scaler = StandardScaler() X_scaled = scaler.fit_transform(X) # ❌ fits on ALL data (leakage) X_train, X_test, y_train, y_test = train_test_split(X_scaled, y) model = LogisticRegression().fit(X_train, y_train) y_pred = model.predict(X_test) # OK by accident
setup added so this can run · defines X, train_test_split, y
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X = _AutoMock('X') def train_test_split(*_a, **_kw): print('-> train_test_split() called') return _AutoMock('train_test_split()') y = _AutoMock('y')
Two faults:
1. Leakage — fit_transform on the whole X means the scaler saw the test set's mean and std. The test set is no longer untouched.
2. Brittle at deployment — to predict on a single new row, you must remember to scale it with the same scaler. If the scaler isn't saved, your production code is silently wrong.
A pipeline removes both bugs by construction.
2. The Pipeline API
A Pipeline is a list of (name, estimator) tuples. All steps except the last must be transformers (have .fit_transform()). The last is the estimator (has .fit()/.predict()).
from sklearn.pipeline import Pipeline from sklearn.preprocessing import StandardScaler from sklearn.linear_model import LogisticRegression pipe = Pipeline([ ("scaler", StandardScaler()), ("clf", LogisticRegression(max_iter=1000)), ])
Now pipe behaves like any other sklearn estimator:
pipe.fit(X_train, y_train) # scales train, fits classifier on scaled train pipe.predict(X_test) # scales test (with TRAIN stats), then predicts pipe.score(X_test, y_test) # ditto
setup added so this can run · defines X_train, y_train, X_test, y_test, pipe
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_train = _AutoMock('X_train') y_train = _AutoMock('y_train') X_test = _AutoMock('X_test') y_test = _AutoMock('y_test') pipe = _AutoMock('pipe')
The scaler is fit only on X_train. When you predict on X_test, the same fitted scaler transforms it. No leakage. No forgotten step.
from sklearn.datasets import make_classification from sklearn.model_selection import train_test_split X, y = make_classification(n_samples=500, n_features=10, random_state=42) X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42) pipe.fit(X_train, y_train) print("Test accuracy:", pipe.score(X_test, y_test)) # → Test accuracy: 0.872
setup added so this can run · defines pipe
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) pipe = _AutoMock('pipe')
3. Mixed Types: ColumnTransformer
Real datasets have numeric and categorical columns that need different preprocessing — scaling for one, one-hot encoding for the other. ColumnTransformer routes each column to the right transformer, then a Pipeline chains the result into the model.
import pandas as pd from sklearn.compose import ColumnTransformer from sklearn.preprocessing import StandardScaler, OneHotEncoder from sklearn.pipeline import Pipeline from sklearn.linear_model import LogisticRegression df = pd.DataFrame({ "age": [25, 32, 47, 51, 28, 39, 60, 22, 45, 33], "income": [45, 62, 88, 95, 41, 70, 110, 38, 79, 58], "country": ["UK", "US", "UK", "DE", "FR", "US", "DE", "FR", "UK", "US"], "bought": [0, 1, 1, 1, 0, 1, 1, 0, 1, 1], }) X = df.drop(columns="bought") y = df["bought"] numeric_features = ["age", "income"] categorical_features = ["country"] preprocessor = ColumnTransformer([ ("num", StandardScaler(), numeric_features), ("cat", OneHotEncoder(handle_unknown="ignore"), categorical_features), ]) pipe = Pipeline([ ("preprocess", preprocessor), ("clf", LogisticRegression(max_iter=1000)), ]) pipe.fit(X, y) print("Train score:", pipe.score(X, y)) # → Train score: 1.0
pipe.fit(X, y) now does all of this in order:
1. Scale age and income with StandardScaler.
2. One-hot encode country into UK/US/DE/FR boolean columns.
3. Concatenate the result.
4. Fit the logistic regression on the combined feature matrix.
pipe.predict(new_row) applies the same steps in the same order, with the same fitted state. Zero glue code.
handle_unknown="ignore" on OneHotEncoder is non-negotiable for production — without it, a new country at inference time crashes the pipeline.
4. Cross-Validation on a Pipeline
Because a pipeline is an estimator, you pass it to cross_val_score like any other model:
from sklearn.model_selection import cross_val_score scores = cross_val_score(pipe, X, y, cv=5) print(f"CV mean: {scores.mean():.3f} ± {scores.std():.3f}") # → CV mean: 0.860 ± 0.045
setup added so this can run · defines pipe, X, y
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) pipe = _AutoMock('pipe') X = _AutoMock('X') y = _AutoMock('y')
Critically: at each of the 5 folds, the pipeline re-fits the scaler on just that fold's training data. Cross-validation across a pipeline is leakage-free by construction. Cross-validation across a manually-scaled dataset is leakage-prone.
This alone is reason enough to wrap your work in a pipeline from line one.
5. Hyperparameter Tuning with GridSearchCV
Pipelines compose with GridSearchCV. The trick is the double-underscore parameter syntax: stepname__paramname.
from sklearn.model_selection import GridSearchCV param_grid = { "clf__C": [0.01, 0.1, 1, 10], # logistic regression's regularisation "clf__solver": ["liblinear", "lbfgs"], } search = GridSearchCV(pipe, param_grid, cv=5, scoring="accuracy") search.fit(X, y) print("Best params:", search.best_params_) print("Best CV: ", round(search.best_score_, 3)) # → Best params: {'clf__C': 1, 'clf__solver': 'lbfgs'} # → Best CV: 0.870
setup added so this can run · defines pipe, X, y
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) pipe = _AutoMock('pipe') X = _AutoMock('X') y = _AutoMock('y')
clf__C reads as "the C parameter of the step named clf". Same pattern for any nested parameter: preprocess__num__with_mean, preprocess__cat__drop. Once you know the rule you can tune anything in the pipeline.
search.best_estimator_ is the refit-on-everything pipeline ready to go. search itself behaves like an estimator too — search.predict(X_new) works directly.
6. Saving a Trained Pipeline
The win you've been building toward: the whole fitted pipeline — preprocessing, model, everything — pickles as one object.
import joblib joblib.dump(pipe, "model.joblib") # Months later, in a totally different script or service: loaded = joblib.load("model.joblib") loaded.predict(X_new) # applies the same scaling, encoding, model — exactly
setup added so this can run · defines pipe, X_new
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) pipe = _AutoMock('pipe') X_new = _AutoMock('X_new')
joblib is the sklearn-recommended serialiser. It handles NumPy arrays efficiently and is a drop-in for pickle for ML objects.
Versioning gotcha: pickled models are tied to the library versions you trained with. If your production environment has scikit-learn 1.2 and you trained on 1.5, expect warnings or outright failures. Pin your versions in requirements.txt and ideally store the version alongside the model file.
Forward link: shipping this model behind an API is covered in deployment. Pipelines are why deployment is easy — you ship one object, not a tangle of preprocessing scripts.
7. The End-to-End Script
Here's the minimum viable real-world ML script. Memorise the shape.
import joblib import pandas as pd from sklearn.model_selection import train_test_split, cross_val_score from sklearn.compose import ColumnTransformer from sklearn.preprocessing import StandardScaler, OneHotEncoder from sklearn.pipeline import Pipeline from sklearn.ensemble import RandomForestClassifier # 1. Load df = pd.read_csv("data.csv") X = df.drop(columns="target") y = df["target"] # 2. Split X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=42) # 3. Pipeline numeric_cols = X.select_dtypes("number").columns.tolist() categorical_cols = X.select_dtypes(exclude="number").columns.tolist() pre = ColumnTransformer([ ("num", StandardScaler(), numeric_cols), ("cat", OneHotEncoder(handle_unknown="ignore"), categorical_cols), ]) pipe = Pipeline([ ("pre", pre), ("clf", RandomForestClassifier(n_estimators=200, random_state=42)), ]) # 4. Validate cv = cross_val_score(pipe, X_train, y_train, cv=5) print(f"CV: {cv.mean():.3f} ± {cv.std():.3f}") # 5. Fit on full training set, score on hold-out pipe.fit(X_train, y_train) print(f"Test: {pipe.score(X_test, y_test):.3f}") # 6. Save joblib.dump(pipe, "model.joblib")
Load CSV → split → pipeline → cross-validate → fit → score → save. Six steps. This script ships to production with maybe 10 more lines (a FastAPI wrapper). You'll write some version of this every week for the rest of your ML career.
Common Mistakes
- Scaling before the split. As shown in Section 1 —
scaler.fit_transform(X)then split. The test set's statistics leaked. Always scale inside the pipeline. - Manually re-doing transforms at inference. "I'll just one-hot the new row in my API handler." No — load the pipeline, call
.predict(), done. Manual reimplementation drifts from training and fails silently. - Not pickling the pipeline as one unit. Saving model and encoder separately, then forgetting to apply both. Make the pipeline the only thing you serialise.
- Forgetting
handle_unknown="ignore"onOneHotEncoder. First unfamiliar category at inference → crash. - Not pinning library versions. Your
model.joblibbecomes unloadable after a routinepip upgrade. Alwayspip freeze > requirements.txtfor any model that's leaving your laptop.
🎯 Your Turn — Build a Complete Pipeline
Build a full mixed-type pipeline on inline data, cross-validate it, and save the trained version.
Inline data (one numeric, one categorical column, binary target):
import pandas as pd df = pd.DataFrame({ "age": [22, 25, 31, 35, 41, 47, 52, 58, 29, 33, 44, 50, 27, 38, 55], "country": ["UK","US","UK","DE","FR","US","DE","UK","FR","US","DE","UK","FR","US","DE"], "bought": [0, 0, 1, 1, 1, 1, 1, 1, 0, 0, 1, 1, 0, 1, 1], })
Your task:
1. Build a Pipeline with a ColumnTransformer (scale age, one-hot country) followed by LogisticRegression(max_iter=1000).
2. Run 5-fold cross_val_score on the whole df (no train/test split needed — small data).
3. Fit the pipeline on all of df and save it to "pipeline.joblib".
4. Return round(cv_scores.mean(), 3).
import pandas as pd import joblib from sklearn.compose import ColumnTransformer from sklearn.preprocessing import StandardScaler, OneHotEncoder from sklearn.pipeline import Pipeline from sklearn.linear_model import LogisticRegression from sklearn.model_selection import cross_val_score def build_and_save_pipeline(): df = pd.DataFrame({ "age": [22, 25, 31, 35, 41, 47, 52, 58, 29, 33, 44, 50, 27, 38, 55], "country": ["UK","US","UK","DE","FR","US","DE","UK","FR","US","DE","UK","FR","US","DE"], "bought": [0, 0, 1, 1, 1, 1, 1, 1, 0, 0, 1, 1, 0, 1, 1], }) X = df.drop(columns="bought") y = df["bought"] # TODO 1: build a ColumnTransformer # ("num", StandardScaler(), ["age"]) # ("cat", OneHotEncoder(handle_unknown="ignore"), ["country"]) # TODO 2: build a Pipeline ("pre", preprocessor), ("clf", LogisticRegression(max_iter=1000)) # TODO 3: cv = cross_val_score(pipe, X, y, cv=5) # TODO 4: pipe.fit(X, y); joblib.dump(pipe, "pipeline.joblib") # TODO 5: return round(cv.mean(), 3) pass print(build_and_save_pipeline())
Hint 1 — ColumnTransformer takes a list of tuples
Each tuple is(name, transformer, columns). columns is a list of column names, even if there's only one. ["age"] not "age".
Hint 2 — save right before returning
pipe.fit(X, y) mutates pipe in place and also returns it. After fitting, joblib.dump(pipe, "pipeline.joblib"). Then return the rounded CV mean.
Show full solution
import pandas as pd import joblib from sklearn.compose import ColumnTransformer from sklearn.preprocessing import StandardScaler, OneHotEncoder from sklearn.pipeline import Pipeline from sklearn.linear_model import LogisticRegression from sklearn.model_selection import cross_val_score def build_and_save_pipeline(): df = pd.DataFrame({ "age": [22, 25, 31, 35, 41, 47, 52, 58, 29, 33, 44, 50, 27, 38, 55], "country": ["UK","US","UK","DE","FR","US","DE","UK","FR","US","DE","UK","FR","US","DE"], "bought": [0, 0, 1, 1, 1, 1, 1, 1, 0, 0, 1, 1, 0, 1, 1], }) X = df.drop(columns="bought") y = df["bought"] preprocessor = ColumnTransformer([ ("num", StandardScaler(), ["age"]), ("cat", OneHotEncoder(handle_unknown="ignore"), ["country"]), ]) pipe = Pipeline([ ("pre", preprocessor), ("clf", LogisticRegression(max_iter=1000)), ]) cv = cross_val_score(pipe, X, y, cv=5) pipe.fit(X, y) joblib.dump(pipe, "pipeline.joblib") return round(cv.mean(), 3) print(build_and_save_pipeline()) # → 0.733 (small dataset — your CV score will be noisy; anything in this range is fine)
You now have a pipeline.joblib on disk. To use it from any other script:
import joblib, pandas as pd loaded = joblib.load("pipeline.joblib") new_customer = pd.DataFrame([{"age": 40, "country": "UK"}]) print(loaded.predict(new_customer)) # [1] print(loaded.predict_proba(new_customer)) # [[0.21 0.79]]
The loaded pipeline scales the age, one-hot-encodes the country with the exact categories it learned at training, and runs the classifier — all behind one .predict() call. That's the entire payoff of pipelines.
What You Learned
- A
Pipelinechains preprocessing and a model into one estimator. Same.fit()/.predict()/.score()API. - Pipelines eliminate leakage by re-fitting transformers on each fold's training data, and eliminate brittleness by applying the same transforms at inference.
ColumnTransformerroutes different columns to different transformers — scaling for numeric, one-hot for categorical.- Hyperparameter tuning uses the
step__paramdouble-underscore syntax insideGridSearchCV. joblib.dumpserialises the whole fitted pipeline as one file;joblib.loadbrings it back.- The end-to-end script is six steps: load → split → pipeline → cross-validate → fit → save.
You now have everything you need to build a working ML feature: data in, predictions out, model on disk, scoreable, reproducible.
Where to next:
- Trees and ensembles — go deeper on random forests and gradient boosting.
- Deployment — wrap your
.joblibin a FastAPI endpoint. - Deep learning intro — when tabular ML stops being enough.
- LLMs and the OpenAI API — the other branch of the stack.
Practice this
on practicepython.inShort exercises that run in your browser and tell you what your code actually did, not just whether a test passed.