Neural Networks: Architectures and Theory
1 · The lesson
readYou already know that a neural network is layers of weighted sums followed by non-linearities, trained by gradient descent. That's the tourist view. This lesson goes underneath: the linear-algebra notation that papers actually use, why your network with tanh activations refused to train, why the Adam optimiser eats the world, and why a 152-layer ResNet is even possible.
Everything here is grounded in one fact: training a deep network is an optimisation problem in millions of parameters, with a non-convex loss surface, fitted by a noisy first-order method. Every trick in this lesson — activation choice, initialisation, normalisation, residuals, optimisers, schedules — exists because that fact has consequences, and the consequences will bite you.
1. The MLP in Formal Notation
A multi-layer perceptron (MLP) is the bedrock architecture. Layer $\ell$ takes input $\mathbf{a}^{(\ell-1)} \in \mathbb{R}^{n_{\ell-1}}$ and produces:
$$
\mathbf{z}^{(\ell)} = W^{(\ell)} \mathbf{a}^{(\ell-1)} + \mathbf{b}^{(\ell)}, \qquad \mathbf{a}^{(\ell)} = \sigma\!\left(\mathbf{z}^{(\ell)}\right)
$$
- $W^{(\ell)} \in \mathbb{R}^{n_\ell \times n_{\ell-1}}$ — the weight matrix
- $\mathbf{b}^{(\ell)} \in \mathbb{R}^{n_\ell}$ — the bias
- $\sigma$ — the activation function applied element-wise
- $\mathbf{z}^{(\ell)}$ — the pre-activation (logits, before $\sigma$)
- $\mathbf{a}^{(\ell)}$ — the activation (after $\sigma$)
The full forward pass for an $L$-layer network is the function composition:
$$
f(\mathbf{x}) = \sigma_L \circ W_L \circ \sigma_{L-1} \circ W_{L-1} \circ \cdots \circ \sigma_1 \circ W_1 (\mathbf{x})
$$
Without the $\sigma$'s, the whole thing collapses to a single linear map $W_L W_{L-1} \cdots W_1 \mathbf{x}$ — no matter how many layers you stack. The non-linearity is what gives the network its expressive power. A network of any depth with only linear activations is exactly equivalent to one linear layer.
In code, in pure NumPy, this is unambiguous:
import numpy as np def forward(x, params): a = x for W, b, sigma in params: z = W @ a + b a = sigma(z) return a
That's it. Everything else in this lesson is about making that loop trainable.
2. Activation Functions — Picking the Right Non-Linearity
The activation function determines the gradient signal that flows back through the network. Pick badly and your gradients vanish, explode, or get stuck at zero forever.
Sigmoid — historically dominant, modernly avoided
$$
\sigma(z) = \frac{1}{1+e^{-z}}, \qquad \sigma'(z) = \sigma(z)(1-\sigma(z))
$$
The derivative peaks at $0.25$ when $z = 0$ and decays to zero on both sides. Stack ten sigmoid layers and the gradient is multiplied by at most $0.25^{10} \approx 10^{-6}$ — the vanishing gradient problem. Sigmoid still has a home in the output layer of a binary classifier (it maps $\mathbb{R} \to (0,1)$ cleanly), but in hidden layers it's a museum piece.
Tanh — sigmoid's zero-centred cousin
$$
\tanh(z) = \frac{e^z - e^{-z}}{e^z + e^{-z}}, \qquad \tanh'(z) = 1 - \tanh^2(z)
$$
Zero-centred output $(-1, 1)$ — slightly better for optimisation than sigmoid. Still saturates on both sides. Used in older RNNs and in the gates of LSTMs.
ReLU — the workhorse since 2012
$$
\text{ReLU}(z) = \max(0, z), \qquad \text{ReLU}'(z) = \begin{cases}1 & z > 0 \\ 0 & z \le 0\end{cases}
$$
Cheap, no saturation on the positive side, gradient is exactly $1$ for active units — gradients flow through deep stacks without decay. ReLU is responsible for making networks with 20+ layers trainable.
The dying-ReLU problem: a unit that gets pushed into $z < 0$ by an unlucky update has gradient $0$ forever — no future gradient will ever reach it. Dead units are permanently dead. With a high learning rate, you can kill 40% of your network in the first epoch and never recover.
Leaky ReLU — keep the dead units alive
$$
\text{LReLU}(z) = \begin{cases}z & z > 0 \\ \alpha z & z \le 0\end{cases}, \qquad \alpha = 0.01 \text{ (typical)}
$$
A small negative slope means dead units still get a gradient. Cheap insurance against the dying-ReLU problem.
GELU — the transformer default
$$
\text{GELU}(z) = z \cdot \Phi(z)
$$
where $\Phi$ is the standard normal CDF. Smooth, non-monotonic, weights inputs by their value rather than gating hard at zero. Used in BERT, GPT, ViT — pretty much every modern transformer. Slightly more expensive than ReLU; in practice it usually wins on transformer language models.
Swish / SiLU — used in EfficientNet, Llama
$$
\text{Swish}(z) = z \cdot \sigma(\beta z)
$$
Self-gated, smooth, similar performance to GELU. nn.SiLU() in PyTorch, keras.activations.swish in Keras.
When to use what — the cheat sheet
| Activation | Use when |
|---|---|
| ReLU | Default for CNNs and most MLPs. Cheap, fast, works. |
| Leaky ReLU / PReLU | You're seeing dead units, or training very deep without BN. |
| GELU | Transformers. Now the default for LLMs and ViT. |
| Swish/SiLU | EfficientNet-style CNNs, modern Llama-family LLMs. |
| Tanh | LSTM/GRU gates. Almost nowhere else. |
| Sigmoid | Output layer of binary classifier. Never hidden. |
| Softmax | Output layer of multi-class classifier. |
3. Weight Initialisation — Why Random Isn't Enough
If every weight starts at the same value, every hidden unit computes the same thing — symmetry. Gradient descent updates them identically. You effectively have a one-unit-per-layer network. So weights must be random.
But how random matters enormously. Consider a layer with $n_{\text{in}}$ inputs. The pre-activation variance is:
$$
\operatorname{Var}(z) = n_{\text{in}} \cdot \operatorname{Var}(W) \cdot \operatorname{Var}(a)
$$
If $\operatorname{Var}(W)$ is too large, $\operatorname{Var}(z)$ grows by $n_{\text{in}}$ at every layer — pre-activations explode, gradients explode. If $\operatorname{Var}(W)$ is too small, signal vanishes layer by layer.
Goal: pick $\operatorname{Var}(W)$ so that signal variance is preserved layer to layer.
Xavier / Glorot initialisation (2010)
For symmetric activations (tanh, sigmoid):
$$
W \sim \mathcal{N}\!\left(0, \frac{2}{n_{\text{in}} + n_{\text{out}}}\right) \quad \text{or} \quad W \sim \mathcal{U}\!\left(-\sqrt{\tfrac{6}{n_{\text{in}}+n_{\text{out}}}}, \sqrt{\tfrac{6}{n_{\text{in}}+n_{\text{out}}}}\right)
$$
Keeps variance constant in both directions — forward (activations) and backward (gradients).
He initialisation (2015)
For ReLU and its variants — ReLU zeros out half the inputs, so you need double the variance to compensate:
$$
W \sim \mathcal{N}\!\left(0, \frac{2}{n_{\text{in}}}\right)
$$
Always use He init when your activations are ReLU/LReLU/GELU/Swish. Keras: kernel_initializer="he_normal". PyTorch: nn.init.kaiming_normal_.
# Keras — pick init to match activation from tensorflow.keras import layers, initializers layers.Dense(128, activation="relu", kernel_initializer=initializers.HeNormal()) layers.Dense(64, activation="tanh", kernel_initializer=initializers.GlorotNormal())
Mismatch the init and the activation and your loss curve will look weird from epoch 1 — flat, NaN, or stuck. Match them and the first few epochs converge cleanly.
4. Backpropagation — A Worked Example
Let's do one full backward pass by hand. Two-layer network, scalar input, scalar output, MSE loss. No batching, no broadcasting, just chain rule.
Forward:
$$
z_1 = w_1 x + b_1, \quad a_1 = \text{ReLU}(z_1), \quad z_2 = w_2 a_1 + b_2, \quad \hat{y} = z_2
$$
$$
\mathcal{L} = \tfrac{1}{2}(\hat{y} - y)^2
$$
Concrete numbers: $x = 2$, $y = 1$, $w_1 = 0.5$, $b_1 = 0$, $w_2 = -1$, $b_2 = 0.3$.
- $z_1 = 0.5 \cdot 2 + 0 = 1.0$
- $a_1 = \text{ReLU}(1.0) = 1.0$
- $z_2 = -1 \cdot 1.0 + 0.3 = -0.7$
- $\hat{y} = -0.7$
- $\mathcal{L} = \tfrac{1}{2}(-0.7 - 1)^2 = \tfrac{1}{2}(2.89) = 1.445$
Backward — work from the output back, multiplying derivatives (chain rule):
$$
\frac{\partial \mathcal{L}}{\partial \hat{y}} = \hat{y} - y = -1.7
$$
$$
\frac{\partial \mathcal{L}}{\partial z_2} = \frac{\partial \mathcal{L}}{\partial \hat{y}} \cdot \frac{\partial \hat{y}}{\partial z_2} = -1.7 \cdot 1 = -1.7
$$
$$
\frac{\partial \mathcal{L}}{\partial w_2} = \frac{\partial \mathcal{L}}{\partial z_2} \cdot a_1 = -1.7 \cdot 1.0 = -1.7
$$
$$
\frac{\partial \mathcal{L}}{\partial b_2} = \frac{\partial \mathcal{L}}{\partial z_2} = -1.7
$$
$$
\frac{\partial \mathcal{L}}{\partial a_1} = \frac{\partial \mathcal{L}}{\partial z_2} \cdot w_2 = -1.7 \cdot -1 = 1.7
$$
$$
\frac{\partial \mathcal{L}}{\partial z_1} = \frac{\partial \mathcal{L}}{\partial a_1} \cdot \text{ReLU}'(z_1) = 1.7 \cdot 1 = 1.7 \quad (\text{since } z_1 > 0)
$$
$$
\frac{\partial \mathcal{L}}{\partial w_1} = \frac{\partial \mathcal{L}}{\partial z_1} \cdot x = 1.7 \cdot 2 = 3.4
$$
$$
\frac{\partial \mathcal{L}}{\partial b_1} = \frac{\partial \mathcal{L}}{\partial z_1} = 1.7
$$
Gradient descent step with learning rate $\eta = 0.1$:
- $w_1 \leftarrow 0.5 - 0.1 \cdot 3.4 = 0.16$
- $w_2 \leftarrow -1 - 0.1 \cdot (-1.7) = -0.83$
- $b_1 \leftarrow 0 - 0.1 \cdot 1.7 = -0.17$
- $b_2 \leftarrow 0.3 - 0.1 \cdot (-1.7) = 0.47$
That is one training step. Vector-valued inputs, batched data, and matrix weights generalise this exact pattern — every gradient is a product of partials along the path from $\mathcal{L}$ back to the parameter. Autograd in TensorFlow/PyTorch builds the computation graph and runs this multiplication for you, but it's the same chain rule on each edge.
If $z_1$ had been negative, $\text{ReLU}'(z_1) = 0$ would kill the entire upstream gradient at that node — that's why dying ReLU is fatal.
5. Loss Functions — More Than MSE
The loss defines what "correct" means. Pick wrong, and the optimiser will happily converge to something useless.
| Loss | Formula | When |
|---|---|---|
| MSE | $\frac{1}{n}\sum (y-\hat{y})^2$ | Regression, Gaussian noise, no outliers |
| MAE / L1 | $\frac{1}{n}\sum \lvert y-\hat{y} \rvert$ | Regression with outliers — less sensitive |
| Huber | quadratic near zero, linear far away | Regression that wants both — robust default |
| Binary cross-entropy | $-[y \log\hat{y} + (1-y)\log(1-\hat{y})]$ | Binary classification with sigmoid output |
| Categorical cross-entropy | $-\sum_k y_k \log \hat{y}_k$ | Multi-class with softmax output (one-hot $y$) |
| Sparse cat. cross-entropy | same, but $y$ is integer label | Same as above, cheaper when classes are huge |
| Focal loss | $-(1-\hat{y})^\gamma \log\hat{y}$ | Heavy class imbalance — down-weights easy examples |
| Contrastive / triplet | based on distances in embedding space | Metric learning (face recognition, embeddings) |
The loss changes the geometry of the problem. MSE on classification logits is convex in the wrong way — it under-penalises confident wrong predictions. Cross-entropy on regression is undefined for negative residuals. Pick the loss for the task before you do anything else.
Focal loss deserves a special note: in detection, 99% of anchor boxes are background. Plain cross-entropy gets dominated by the easy negatives. Focal loss adds a $(1-\hat{y})^\gamma$ factor that suppresses confident-correct examples and forces the network to focus on hard ones.
6. Optimisers — The Update Rules
Vanilla SGD updates parameters with $\theta \leftarrow \theta - \eta \nabla_\theta \mathcal{L}$. It works, but it's slow and gets stuck in ravines and saddle points. Every modern optimiser is SGD plus tricks.
SGD with momentum
Accumulates an exponentially weighted velocity:
$$
v_t = \mu v_{t-1} + \nabla_\theta \mathcal{L}, \qquad \theta \leftarrow \theta - \eta v_t
$$
$\mu = 0.9$ is the canonical choice. Speeds up consistent directions, damps oscillation in noisy ones.
Nesterov momentum
Look-ahead: compute gradient at $\theta - \eta \mu v_{t-1}$ rather than $\theta$. Tighter convergence in convex problems; in deep learning the gap is small.
RMSprop
Per-parameter learning rate, scaled by recent gradient magnitude:
$$
s_t = \beta s_{t-1} + (1-\beta)\, g_t^2, \qquad \theta \leftarrow \theta - \frac{\eta}{\sqrt{s_t}+\epsilon}\, g_t
$$
Where $g_t = \nabla_\theta \mathcal{L}$. Parameters with consistently large gradients get a smaller effective rate. Dominant choice for RNNs before Adam.
Adam — adaptive moment estimation
Momentum + RMSprop combined:
$$
m_t = \beta_1 m_{t-1} + (1-\beta_1) g_t \quad (\text{first moment})
$$
$$
v_t = \beta_2 v_{t-1} + (1-\beta_2) g_t^2 \quad (\text{second moment})
$$
$$
\hat{m}_t = \frac{m_t}{1-\beta_1^t}, \quad \hat{v}_t = \frac{v_t}{1-\beta_2^t} \quad (\text{bias correction})
$$
$$
\theta \leftarrow \theta - \frac{\eta}{\sqrt{\hat{v}_t}+\epsilon}\, \hat{m}_t
$$
Defaults: $\beta_1=0.9$, $\beta_2=0.999$, $\epsilon=10^{-8}$, $\eta=10^{-3}$. Adam works on almost everything out of the box — it's the right starting point unless you have a reason not to.
AdamW — decoupled weight decay
Original Adam applies L2 regularisation by adding $\lambda \theta$ to the gradient, then dividing by $\sqrt{\hat{v}}$ — meaning weight decay gets scaled by the per-parameter learning rate. Loshchilov & Hutter showed this breaks the intended regularisation effect. AdamW decouples:
$$
\theta \leftarrow \theta - \eta \left(\frac{\hat{m}_t}{\sqrt{\hat{v}_t}+\epsilon} + \lambda \theta\right)
$$
The $\lambda \theta$ term is applied separately, not coupled to the adaptive scaling. AdamW is the de facto optimiser for transformers and most modern training. Keras: keras.optimizers.AdamW(weight_decay=0.01).
7. Learning Rate Schedules
A single fixed learning rate is rarely right for the whole of training. You want fast early to escape saddle points, slow late to settle into a sharp minimum.
- Step decay: drop $\eta$ by a factor every $k$ epochs. Simple, manual.
- Exponential decay: $\eta_t = \eta_0 \gamma^t$. Smoother than step decay.
- Cosine annealing: $\eta_t = \eta_{\min} + \tfrac{1}{2}(\eta_{\max} - \eta_{\min})(1 + \cos(\pi t / T))$. Default for modern training. Smooth descent to near-zero.
- Warmup: linearly ramp $\eta$ from $0$ to $\eta_0$ over the first $w$ steps. Critical for transformers — large models with adaptive optimisers fail without it (early steps have unreliable variance estimates).
- OneCycle (Smith, 2017): warmup then cosine decay, plus inverse momentum schedule. Aggressive but extremely effective for CNNs.
import tensorflow as tf schedule = tf.keras.optimizers.schedules.CosineDecay( initial_learning_rate=1e-3, decay_steps=10_000, alpha=0.01, # final LR = alpha * initial ) opt = tf.keras.optimizers.AdamW(learning_rate=schedule, weight_decay=0.01)
The schedule lives inside the optimiser. Pass learning_rate=schedule and Keras handles the per-step update.
8. Batch Normalisation — What It Actually Does
For each mini-batch and each feature, BatchNorm normalises:
$$
\hat{z} = \frac{z - \mu_{\text{batch}}}{\sqrt{\sigma^2_{\text{batch}} + \epsilon}}, \qquad \text{out} = \gamma \hat{z} + \beta
$$
$\gamma, \beta$ are learned per-feature scale and shift. At test time, $\mu$ and $\sigma$ come from running averages tracked during training.
The original (2015) paper claimed BN works by reducing "internal covariate shift". That explanation is mostly wrong — a 2018 paper by Santurkar et al. showed BN works by smoothing the loss landscape, making gradients better-behaved. Whatever the exact mechanism, the empirical effect is: BN lets you use higher learning rates, train deeper networks, and acts as a mild regulariser.
Critical gotcha: BN has different behaviour at train vs eval time. Forget to set model.trainable = False (Keras) or .eval() (PyTorch) at inference and your running stats keep updating with whatever distribution you happen to feed in — silently breaking the model.
9. Layer / Group / Instance Normalisation
BatchNorm breaks down when the batch is small (batch=1 in RL, batch=2 in 3D vision). Alternatives normalise across different axes:
| Norm | Normalises over |
|---|---|
| BatchNorm | Each feature across the batch dimension |
| LayerNorm | All features within a single sample |
| InstanceNorm | Each feature within a single sample (per spatial position) — style transfer |
| GroupNorm | Each group of features within a single sample — middle ground |
LayerNorm is the norm of choice for transformers: it doesn't depend on batch size at all, which makes training stable across hardware and sequence lengths. Every transformer block ends in LayerNorm. GroupNorm is the choice for small-batch CV — common in detection and segmentation where batch=2 is normal.
10. Residual Connections — Why We Can Go Deep
A residual block computes $y = F(x) + x$ rather than $y = F(x)$. The skip lets gradient flow directly through:
$$
\frac{\partial \mathcal{L}}{\partial x} = \frac{\partial \mathcal{L}}{\partial y} \cdot \left(\frac{\partial F}{\partial x} + I\right)
$$
The $+I$ term means the gradient always has a "shortcut" path — even if $\partial F/\partial x$ vanishes, the gradient still propagates back through the identity. ResNet's contribution wasn't deeper networks per se — it was the architecture that made deep training stable. Before ResNet, networks past ~20 layers got worse with depth (not from overfitting — from optimisation failure). After ResNet, 1000-layer networks were trainable.
Every modern architecture uses residuals: transformers, U-Nets, EfficientNet, ConvNeXt. If you build something deeper than ~10 layers, add residual connections. It's not optional.
11. The Bias-Variance Trade-off, DL Edition
Classical statistical learning: too few parameters → high bias (underfitting), too many → high variance (overfitting). The optimal model is in the middle.
Deep learning broke this. Modern networks are massively over-parameterised — billions of parameters, tens of thousands of training examples — and still generalise. The classical story would predict catastrophic overfitting. It doesn't happen.
Double descent (Belkin et al., 2019): plot test error against model size. As you increase model size, error first rises (the classical U-curve), then peaks at the interpolation threshold (where the model can exactly fit the training set), then drops again as the model gets even bigger. The right side of the curve is where modern DL lives.
The mechanism is implicit regularisation — SGD biases toward "simple" solutions in the over-parameterised regime, plus features like skip connections, BN, and weight decay add explicit regularisation. The classical bias-variance picture isn't wrong, but it's incomplete.
Practical takeaway: don't be afraid to use a model that "should" overfit. With proper regularisation, the second descent is real and consistent.
Common Mistakes
1. Forgetting normalisation between layers
A 12-layer MLP with ReLU and no BatchNorm/LayerNorm will train, but slowly and with high variance across seeds. Add a normalisation layer after every Dense/Conv (or use a normalisation scheme designed into the block) — it's free regularisation and a 2-5× speedup.
2. Same LR for all layers when fine-tuning
When fine-tuning a pretrained model on a new task, the head needs a normal LR but the backbone needs a tiny LR (the pretrained features are mostly correct). Use layer-wise LR decay or a small global LR (1e-5) and only the head at 1e-3. Lumping them together either destroys pretrained features or undertrains the head.
3. Treating losses as drop-in replacements
Swapping MSE for cross-entropy isn't just a formula change — it's a different optimisation landscape. Cross-entropy with logits is much steeper near wrong predictions; the same learning rate that worked for MSE will explode with CE. Always retune LR when changing loss.
4. Init mismatch
Dense(128, activation="relu") with default Glorot init will train, but slower than Dense(128, activation="relu", kernel_initializer="he_normal"). With deep networks, the gap is large enough to matter — sometimes the difference between converging and not.
5. Ignoring the warmup
Train a transformer or large vision model with constant LR from step 0 and you'll see early loss spikes or NaNs. Adam's variance estimates are unreliable for the first few hundred steps. A 1000-step linear warmup costs you almost nothing and prevents the failure mode.
🎯 Your Turn — A 2-Layer MLP Forward Pass in Pure NumPy
Implement a forward pass for a 2-layer classifier: input → Dense(hidden, ReLU) → Dense(num_classes, softmax). No Keras, no autograd — just NumPy and the equations from Section 1.
Requirements:
- Accept a batch
Xof shape(batch, in_features). - Weights are passed as a tuple
(W1, b1, W2, b2).
W1 shape: (hidden, in_features), b1 shape: (hidden,)
- W2 shape: (num_classes, hidden), b2 shape: (num_classes,)
- Apply ReLU on the hidden layer pre-activations.
- Apply numerically stable softmax on the output logits.
- Return a probability matrix of shape
(batch, num_classes), rows summing to 1.
import numpy as np def forward(X, params): W1, b1, W2, b2 = params # TODO 1: compute hidden pre-activations z1 = X @ W1.T + b1 # TODO 2: apply ReLU a1 = max(0, z1) # TODO 3: compute output logits z2 = a1 @ W2.T + b2 # TODO 4: numerically stable softmax along axis=1 ... # Sanity check np.random.seed(0) X = np.random.randn(4, 3) W1 = np.random.randn(8, 3) * np.sqrt(2/3) # He init b1 = np.zeros(8) W2 = np.random.randn(2, 8) * np.sqrt(2/8) b2 = np.zeros(2) probs = forward(X, (W1, b1, W2, b2)) print(probs.shape) # (4, 2) print(probs.sum(axis=1)) # [1. 1. 1. 1.]
Hint 1 — Broadcasting the bias
X @ W1.T has shape (batch, hidden). Adding b1 of shape (hidden,) broadcasts correctly — NumPy treats b1 as (1, hidden). Don't reshape or transpose b1; just add.
Hint 2 — Numerically stable softmax
A naive softmaxexp(z) / exp(z).sum() overflows for large logits (e.g. exp(1000) = inf). Subtract the per-row max first: z - z.max(axis=1, keepdims=True). Mathematically identical (softmax is shift-invariant), numerically safe. The keepdims=True matters for broadcasting.
Show full solution
import numpy as np def forward(X, params): """Forward pass: Dense(ReLU) -> Dense(softmax).""" W1, b1, W2, b2 = params # Hidden layer z1 = X @ W1.T + b1 # (batch, hidden) a1 = np.maximum(0, z1) # ReLU, element-wise # Output layer z2 = a1 @ W2.T + b2 # (batch, num_classes) # Numerically stable softmax (subtract per-row max) z2_shift = z2 - z2.max(axis=1, keepdims=True) exp_z2 = np.exp(z2_shift) probs = exp_z2 / exp_z2.sum(axis=1, keepdims=True) return probs np.random.seed(0) X = np.random.randn(4, 3) W1 = np.random.randn(8, 3) * np.sqrt(2/3) b1 = np.zeros(8) W2 = np.random.randn(2, 8) * np.sqrt(2/8) b2 = np.zeros(2) probs = forward(X, (W1, b1, W2, b2)) print(probs.shape) # (4, 2) print(probs.sum(axis=1)) # [1. 1. 1. 1.] — rows sum to 1 print(np.allclose(probs.sum(axis=1), 1.0)) # True
Two design choices worth noting:
W1.Tconvention: storing weights as(out, in)matches Keras/PyTorch's internal layout, where each row is one output unit's weight vector. The transpose in the matmul is the cost of that convention. Some texts use(in, out)and skip the transpose — pick one and be consistent.keepdims=Truein the softmax max and sum: without it, the reduction collapses the axis and broadcasting silently misaligns. This is a subtle bug source — always passkeepdims=Truewhen you'll broadcast the result back.
To make this trainable, you'd add a backward pass: cross-entropy gradient at the output ($\hat{y} - y$ for one-hot labels — the softmax+CE pair has that beautifully simple gradient), backprop through W2, through ReLU's $\{0,1\}$ mask, through W1. Each step is a transpose-matmul plus an element-wise operation. The whole NumPy MLP trainer fits in 40 lines — a worthwhile exercise to write once.
What You Learned
- The MLP is layered $\mathbf{a}^{(\ell)} = \sigma(W^{(\ell)} \mathbf{a}^{(\ell-1)} + \mathbf{b}^{(\ell)})$. Without non-linearity, depth is wasted.
- Activations: ReLU is the default; GELU/Swish for transformers and modern CNNs; sigmoid/softmax only at the output. The dying ReLU problem is real — Leaky ReLU is the cheap fix.
- Initialisation must match the activation: He for ReLU-family, Xavier/Glorot for tanh/sigmoid. Mismatch causes vanishing/exploding signal.
- Backprop is the chain rule applied along the computation graph — one multiply per edge. Autograd automates it; the math doesn't change.
- Loss functions define the geometry of the problem. Cross-entropy for classification, Huber for robust regression, focal for severe class imbalance.
- Adam/AdamW is the default optimiser. SGD+momentum still wins on vision sometimes. Warmup + cosine decay is the modern schedule.
- BatchNorm smooths the loss landscape (despite the original "covariate shift" framing). LayerNorm for transformers, GroupNorm for small-batch CV.
- Residual connections are what made deep networks trainable. Use them past ~10 layers; non-negotiable in modern designs.
- Double descent means over-parameterisation can help generalisation — the classical bias-variance picture is incomplete in the DL regime.
Next: TensorFlow & Keras: Production Engine — the framework that turns these primitives into trained, deployed, and scaled models.