PythonMastery
beginner 24 min read · lesson 5 of 7 in Deep Learning Fundamentals

Convolutional Neural Networks

1 · The lesson

read

A 224×224 RGB image is 150,528 numbers. A fully-connected (Dense) network with a single 1,000-neuron hidden layer would need 150 million weights just for that first layer — most of them learning the same edge detector at slightly different pixel positions. That's wasteful, slow, and a recipe for overfit.

Convolutional Neural Networks fix this with two ideas: local receptive fields (each filter looks at a small patch) and weight sharing (the same filter slides across the whole image). The result is a model that is dramatically smaller, translation-invariant by construction, and stupidly good at vision.

Run in Colab (free GPU strongly recommended) or pip install tensorflow. Outputs shown inline.


1. The Convolution — A Sliding Filter

Imagine a small 3×3 grid of weights — a filter or kernel. Slide it across the image one pixel at a time. At each position, compute a weighted sum of the 9 pixels under the filter. The output is a new grid — a feature map — slightly smaller than the input.

python
Image (5×5)        Filter (3×3)        Output (3×3)
[1 1 1 0 0]        [1 0 -1]            [.  .  .]
[0 1 1 1 0]        [1 0 -1]   = slide  [.  .  .]
[0 0 1 1 1]        [1 0 -1]            [.  .  .]
[0 0 1 1 0]
[0 1 1 0 0]

The same 9 weights are used at every position. That's the weight sharing that makes CNNs efficient — a CNN with 10 filters of size 3×3 has only 9·10 = 90 weights for its first layer, regardless of image size.

A filter learns one pattern: maybe a vertical edge, a 45° edge, a red blob. A convolutional layer is just a stack of filters, each learning a different pattern.


2. The Hierarchy Visualised

What CNNs actually learn, layer by layer:

  • Layer 1 — oriented edges and colour blobs. The same kinds of features the visual cortex uses.
  • Layer 2 — textures, corners, simple curves.
  • Layer 3 — object parts: wheels, eyes, fur patches.
  • Deep layers — whole-object detectors: "car", "cat", "stop sign".

This hierarchy isn't programmed — it emerges from training on labelled images, because useful features at one layer enable more useful features at the next. The first layer of nearly every image CNN ever trained looks like the same set of Gabor-style edge detectors. It's one of the more beautiful empirical results in machine learning.


3. The Canonical Pattern

python
Conv → Pool → Conv → Pool → ... → Flatten → Dense → output
  • Conv2D — extract features (replaces Dense).
  • Pooling — shrink the feature map, building in some invariance.
  • Repeat — deeper layers see larger effective patches.
  • Flatten — turn the 3D feature map into a 1D vector.
  • Dense — classifier on top of the learned features.

Almost every image CNN you'll see follows this template. The fancy modern ones (ResNet, EfficientNet) just add skip connections or change how the layers interconnect.


4. A Minimal Keras CNN

python
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense

model = Sequential([
    Conv2D(32, (3, 3), activation='relu', input_shape=(28, 28, 1)),
    MaxPooling2D((2, 2)),
    Conv2D(64, (3, 3), activation='relu'),
    MaxPooling2D((2, 2)),
    Flatten(),
    Dense(10, activation='softmax'),
])

model.compile(optimizer='adam',
              loss='sparse_categorical_crossentropy',
              metrics=['accuracy'])

model.summary()
# Total params: ~34,800 — orders of magnitude smaller than the Dense equivalent

That's a complete MNIST classifier. Read it as: extract 32 features with 3×3 filters, halve the spatial dimensions, extract 64 features at the next scale, halve again, flatten, classify into 10 digits.

Train it on MNIST:

python
from tensorflow.keras.datasets import mnist
import numpy as np

(X_train, y_train), (X_test, y_test) = mnist.load_data()

# Normalise to [0, 1] and add channel dimension
X_train = X_train.astype('float32') / 255.0
X_test  = X_test.astype('float32')  / 255.0
X_train = X_train[..., np.newaxis]      # shape (60000, 28, 28, 1)
X_test  = X_test[...,  np.newaxis]

model.fit(X_train, y_train, epochs=5, batch_size=64, validation_split=0.1)
# → Epoch 5/5  loss: 0.020 - accuracy: 0.993 - val_loss: 0.035 - val_accuracy: 0.990
+ setup added so this can run · defines model
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

model = _AutoMock('model')

Five epochs to 99% on the canonical "hello world" of computer vision. That's the CNN advantage.


5. Conv2D Parameters

python
Conv2D(filters, kernel_size, strides=(1, 1), padding='valid', activation=None)
+ setup added so this can run · defines Conv2D, filters, kernel_size
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def Conv2D(*_a, **_kw):
    print('-> Conv2D() called')
    return _AutoMock('Conv2D()')
filters = _AutoMock('filters')
kernel_size = _AutoMock('kernel_size')
  • filters — how many independent filters this layer learns. 32, 64, 128 are typical.
  • kernel_size — spatial size of each filter. (3, 3) is the modern default. Larger (5×5, 7×7) sometimes used in early layers of older architectures.
  • strides — pixel step between filter positions. (1, 1) keeps full resolution; (2, 2) halves it (an alternative to pooling).
  • padding — what to do at edges:
- 'valid' — no padding. Output shrinks by kernel_size − 1 per dimension. - 'same' — pad with zeros so output keeps the input's spatial size.
  • activation — almost always 'relu'.

A 3×3 conv with padding='same' is a workhorse — same shape in, same shape out, learns local features.


6. Pooling

A pooling layer aggregates a small region of the feature map into one value. Two flavours:

  • MaxPooling2D((2,2)) — keep the strongest signal in each 2×2 region. Default choice.
  • AveragePooling2D((2,2)) — keep the average. Sometimes preferred in modern architectures.

Pooling does two useful things at once:

1. Shrinks the feature map (less compute downstream).
2. Builds small-shift invariance — if the input shifts a pixel or two, the pooled output barely changes.

Most architectures pool with stride 2, halving each spatial dimension.

A modern alternative: a Conv2D with strides=(2, 2) and no pool — does shrinking and feature extraction in one step. Either pattern is fine.


7. Data Shape — Get This Right or Nothing Works

Keras CNNs expect inputs of shape:

python
(batch_size, height, width, channels)
+ setup added so this can run · defines batch_size, height, width, channels
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

batch_size = _AutoMock('batch_size')
height = _AutoMock('height')
width = _AutoMock('width')
channels = _AutoMock('channels')
  • Grayscale (MNIST) — channels = 1.
  • RGB (CIFAR, ImageNet) — channels = 3.
  • RGBA (with alpha) — channels = 4.

If you have (N, H, W) for grayscale, add the channel axis: X = X[..., np.newaxis].

Pixel values must be normalised — divide by 255 to put them in [0, 1], or subtract a mean to get them roughly centred. Raw 0–255 pixels make training unstable.


8. Data Augmentation — Free Data

The single highest-leverage trick for small image datasets. Apply random transformations during training so the model sees a slightly different version of each image each epoch.

The modern Keras way uses preprocessing layers directly in the model:

python
from tensorflow.keras import layers

data_augmentation = keras.Sequential([
    layers.RandomFlip("horizontal"),
    layers.RandomRotation(0.1),         # ±10% of 2π radians
    layers.RandomZoom(0.1),
    layers.RandomContrast(0.1),
])

model = Sequential([
    data_augmentation,                  # active only during training
    Conv2D(32, (3, 3), activation='relu', input_shape=(32, 32, 3)),
    MaxPooling2D((2, 2)),
    # ... rest of the network
])
+ setup added so this can run · defines Sequential, keras, Conv2D, MaxPooling2D
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def Sequential(*_a, **_kw):
    print('-> Sequential() called')
    return _AutoMock('Sequential()')
keras = _AutoMock('keras')
def Conv2D(*_a, **_kw):
    print('-> Conv2D() called')
    return _AutoMock('Conv2D()')
def MaxPooling2D(*_a, **_kw):
    print('-> MaxPooling2D() called')
    return _AutoMock('MaxPooling2D()')

These layers are no-ops at inference time. You don't need to remember to "turn them off".

Common augmentations and what they buy you:

AugmentationBuilds invariance to
RandomFlip(horizontal)left/right pose (NOT for text, NOT for asymmetric objects)
RandomRotationsmall rotations
RandomZoomscale changes
RandomCropframing differences
RandomBrightness/Contrastlighting conditions

Pick the ones that match real variation in your data. Don't flip handwritten digits — 6 becomes nothing meaningful after a horizontal flip.


9. Transfer Learning — Stand on a Giant's Shoulders

You almost never train a CNN from scratch in industry. You take a network pretrained on millions of ImageNet images (ResNet, EfficientNet, MobileNet), strip the final classifier, and bolt on your own small head:

python
from tensorflow.keras.applications import EfficientNetB0

base = EfficientNetB0(weights='imagenet', include_top=False, input_shape=(224, 224, 3))
base.trainable = False                  # freeze the pretrained weights

model = Sequential([
    base,
    layers.GlobalAveragePooling2D(),
    layers.Dense(128, activation='relu'),
    layers.Dense(num_classes, activation='softmax'),
])
+ setup added so this can run · defines Sequential, num_classes, layers
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def Sequential(*_a, **_kw):
    print('-> Sequential() called')
    return _AutoMock('Sequential()')
num_classes = _AutoMock('num_classes')
layers = _AutoMock('layers')

That model can hit 90%+ accuracy on a custom dataset of 1,000 images. Training from scratch would need 100,000+. Full treatment in transfer.


10. The CNN Mental Checklist

Before you model.fit() an image model, verify:

  • Input shape is (H, W, channels) and the first layer's input_shape matches.
  • Pixels are normalised to [0, 1] or [-1, 1].
  • Labels are the right format for the loss (integer for sparse_categorical_crossentropy, one-hot for categorical_crossentropy).
  • Final layer has correct activation: softmax(N) for N-class, sigmoid(1) for binary.
  • For small datasets — augmentation and/or a pretrained backbone are in the mix.

Common Mistakes

  • Wrong input shape. Grayscale (N, 28, 28) will explode when the first Conv2D expects (N, 28, 28, 1). Add np.newaxis.
  • Not normalising pixels. 0–255 inputs make the first layer's gradients huge. Loss often goes to NaN within a few batches. Divide by 255.
  • Tiny dataset + huge CNN from scratch. Will overfit instantly. Use a pretrained backbone and freeze most of it.
  • Augmenting the test set. Augmentation belongs in training only. The Keras layers.RandomXxx preprocessing layers handle this automatically; older ImageDataGenerator requires a separate non-augmented generator for validation.
  • Forgetting Flatten before Dense. Conv2D outputs 3D feature maps; Dense expects 1D. Without Flatten() (or a GlobalAveragePooling2D()) the shapes won't connect.

🎯 Your Turn — MNIST CNN

Build a small CNN for MNIST. Input shape (28, 28, 1). Two Conv→Pool blocks (32 then 64 filters, 3×3 kernels), then Flatten, then Dense(10, softmax). Compile and train for 5 epochs.

Assume X_train, y_train, X_test, y_test are loaded and normalised (see Section 4).

Skeleton:

python
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense

model = Sequential([
    # TODO 1: Conv2D — 32 filters, 3x3, ReLU, input_shape for MNIST
    # TODO 2: MaxPooling2D 2x2
    # TODO 3: Conv2D — 64 filters, 3x3, ReLU
    # TODO 4: MaxPooling2D 2x2
    # TODO 5: Flatten
    # TODO 6: Dense 10 softmax
])

# TODO 7: compile with adam, sparse_categorical_crossentropy, accuracy
# TODO 8: fit for 5 epochs with batch_size=64 and validation_split=0.1
Hint 1 — MNIST shape input_shape=(28, 28, 1) on the first layer only. Grayscale, so channels=1.
Hint 2 — Which loss Labels are integers 0–9, not one-hot. Use sparse_categorical_crossentropy. If you'd one-hot encoded them, you'd use plain categorical_crossentropy.
Show full solution
python
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Conv2D, MaxPooling2D, Flatten, Dense

model = Sequential([
    Conv2D(32, (3, 3), activation='relu', input_shape=(28, 28, 1)),
    MaxPooling2D((2, 2)),
    Conv2D(64, (3, 3), activation='relu'),
    MaxPooling2D((2, 2)),
    Flatten(),
    Dense(10, activation='softmax'),
])

model.compile(optimizer='adam',
              loss='sparse_categorical_crossentropy',
              metrics=['accuracy'])

history = model.fit(X_train, y_train,
                    epochs=5,
                    batch_size=64,
                    validation_split=0.1,
                    verbose=2)
# → Epoch 5/5  loss: 0.021 - accuracy: 0.993 - val_loss: 0.037 - val_accuracy: 0.989

test_loss, test_acc = model.evaluate(X_test, y_test, verbose=0)
print(f"test accuracy: {test_acc:.3f}")
# → test accuracy: 0.990
+ setup added so this can run · defines X_train, y_train, X_test, y_test
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_train = _AutoMock('X_train')
y_train = _AutoMock('y_train')
X_test = _AutoMock('X_test')
y_test = _AutoMock('y_test')

Around 35,000 parameters, five epochs, 99% accuracy. A fully-connected baseline on the same task tops out around 98% with ten times the parameters. That gap is the CNN advantage.


What You Learned

  • A convolution slides a small filter across the image, sharing weights across positions. This makes CNNs efficient and translation-invariant.
  • The canonical pattern: Conv → Pool → Conv → Pool → Flatten → Dense.
  • Pooling shrinks feature maps and adds small-shift invariance.
  • Inputs must be (N, H, W, channels) and normalised to [0, 1].
  • Data augmentation and transfer learning are the two leverage points for real-world image problems.

Next: Working with Text — NLP Basics — the same playbook, but for sequences of tokens.