How Neural Networks Think
1 · The lesson
readA neural network is a function. Numbers go in, numbers come out, and the function adjusts its own internal parameters until the outputs look like what you wanted. That sentence covers ninety-five percent of what's really happening inside the largest model running on a billion-dollar GPU cluster.
Everything else — convolutions, attention, transformers, diffusion — is a shape of that same machinery. So before any framework code, you need a clear mental model of the smallest unit (a neuron), how they stack into layers, how they learn, and where the metaphor with biology politely stops being useful.
This lesson is concept-first. No Keras yet. By the end you should be able to compute a neuron's output on paper.
1. The Smallest Unit — One Neuron
A neuron takes some numbers, multiplies each by a weight, adds a bias, and passes the result through an activation function.
y = activation(w₁·x₁ + w₂·x₂ + ... + wₙ·xₙ + b)That's the whole thing. Two inputs:
y = activation(w₁·x₁ + w₂·x₂ + b)The weights say how much each input matters. The bias shifts the result up or down so the neuron can fire even when all inputs are zero. The activation function adds the non-linearity that lets a stack of these units learn curves, not just straight lines.
Without an activation function, no matter how many neurons you stack, the whole network collapses into one big linear equation. The activation is what makes deep learning deep.
2. Activation Functions — The Four You'll See
Every framework ships dozens; you'll use four 95% of the time.
| Activation | Formula (intuition) | When to use |
|---|---|---|
| ReLU | max(0, x) — pass positives, kill negatives | Default for hidden layers |
| Sigmoid | Squashes to (0, 1) — an S-curve | Binary classification output |
| Tanh | Squashes to (-1, 1) — centred S-curve | RNN hidden states (historical) |
| Softmax | Turns a vector into probabilities that sum to 1 | Multi-class classification output |
ReLU is the modern workhorse. It's almost embarrassingly simple — return the input if positive, otherwise zero — and yet it solved the vanishing gradient problem that crippled sigmoid-based networks in the 1990s. Use ReLU for hidden layers unless you have a specific reason not to.
Sigmoid was historically used everywhere. Now it lives mostly in the final layer of binary classifiers (output ≈ probability the answer is "yes").
Tanh is sigmoid's cousin, but centred at zero. It mostly hangs around in older RNN architectures.
Softmax is special: it operates on a vector of numbers, not a single scalar. It exponentiates each value and divides by the sum, producing a probability distribution. If your model outputs 10 classes, softmax converts the 10 raw scores ("logits") into 10 probabilities summing to 1.
3. Stacking Neurons — Layers
A layer is a group of neurons that all see the same inputs. Each neuron in the layer has its own weights and bias, so each computes a slightly different function of those inputs.
A typical feed-forward network has three kinds of layers:
- Input layer — just a placeholder for your features. No neurons, no learning.
- Hidden layers — the real workers. One, two, twenty, a hundred. Depth is the number of these.
- Output layer — produces the final prediction. Its size and activation depend on the task (1 + sigmoid for binary, N + softmax for N-class, 1 + linear for regression).
Each hidden layer takes the previous layer's outputs as its inputs. Numbers flow forward, layer by layer, getting transformed at every step.
4. Why Depth Helps — Hierarchical Features
Here is the intuition that justifies deep learning:
- Layer 1 learns very simple patterns from raw input — for images, things like edges and colour gradients.
- Layer 2 combines those simple patterns into slightly more complex ones — textures, corners, simple shapes.
- Layer 3 combines those into still-more-complex parts — wheels, eyes, fur.
- Final layers combine parts into whole-object detectors — "car", "cat", "stop sign".
You didn't tell the network what an edge or a wheel is. It discovered those intermediate concepts because they were useful for reducing its error. Each layer abstracts the layer below it. This is what people mean when they say neural networks "learn representations".
The same hierarchy shows up in language models (characters → words → phrases → meaning), audio models (waveforms → phonemes → words), and protein models (atoms → amino acids → folds).
The practical consequence: with images, audio, and text — domains where raw inputs are far from the labels — depth helps enormously. On clean tabular data, where each column is already a meaningful feature, a shallow model or a tree ensemble often wins. Depth pays off when the model needs to invent its own intermediate concepts.
5. The Forward Pass
You feed inputs into the input layer. Each hidden layer computes its outputs from the previous layer's outputs. Eventually the output layer produces a prediction.
input → hidden 1 → hidden 2 → output x h₁ h₂ ŷ
That's the forward pass. It's pure arithmetic — multiply, add, activate, multiply, add, activate, all the way through. No learning happens here. You're just using the current weights to produce a prediction.
Initially the weights are random, so the prediction will be nonsense. A freshly-initialised binary classifier outputs roughly 0.5 for everything. A freshly-initialised 10-class softmax outputs roughly 0.1 for every class. The learning comes next.
6. The Loss Function — Measuring Wrongness
The loss is a number that says how wrong the prediction was compared to the truth. Lower is better. The whole training process is one giant exercise in pushing the loss downward.
You pick a loss function based on your task:
| Task | Loss | Why |
|---|---|---|
| Regression (predict a number) | Mean Squared Error — average of (ŷ − y)² | Penalises big errors quadratically |
| Binary classification | Binary cross-entropy | Measures probability distance |
| Multi-class classification | Categorical cross-entropy | Probability distance over N classes |
For a single prediction with true label y and predicted probability ŷ, binary cross-entropy is:
loss = -(y · log(ŷ) + (1 − y) · log(1 − ŷ))
Don't memorise it. Know that it grows large when the model is confidently wrong (predicts 0.01 when the answer was 1) and shrinks toward zero when the model is confidently right.
7. Backpropagation in Plain English
Once you have the loss, you need to know how to change the weights to reduce it. That's what backpropagation computes.
Imagine the network as a long chain of operations. Backprop walks that chain backwards from the loss, asking at each step: "if I nudged this weight a tiny bit, how much would the loss change?" The answer is the gradient for that weight.
You don't compute backprop by hand — every framework does it automatically (this is called autograd). But you should know what it produces: one gradient per weight, telling you the direction and magnitude to nudge that weight to lower the loss.
8. Gradient Descent — Rolling the Ball Downhill
Picture the loss as a bumpy landscape. Each weight is one dimension; a network with a million weights lives in a million-dimensional landscape. You can't visualise it, but the math doesn't care.
You're standing at a random point (the random initial weights). The gradient tells you which direction is steepest downhill. You take a small step in that direction:
new_weight = old_weight − learning_rate · gradient
The learning rate is the size of your step. Too big, you overshoot the valley. Too small, you take forever to get anywhere. We'll cover tuning it in Training & Improving.
Do this for every weight, with every batch of training data, for many passes through the dataset, and the weights gradually settle into a configuration where the loss is low.
9. The Training Loop in One Paragraph
For each batch of data: run the forward pass to get predictions, compute the loss against the true labels, run backpropagation to get the gradient of the loss with respect to every weight, then update every weight a small step in the downhill direction. Repeat until the loss stops dropping. That single paragraph is the engine of every neural network ever trained.
Some vocabulary that pins this down:
- A batch is a small group of training examples (commonly 32 or 64) processed together.
- One full pass through the training set is an epoch. You usually train for tens to hundreds of epochs.
- The number of weight updates is
(training_set_size / batch_size) · epochs. For a million-example dataset, batch 32, 10 epochs: 312,500 updates.
Each update is tiny. Convergence comes from doing huge numbers of them.
10. The Universal Function Approximator
A celebrated 1989 theorem says a neural network with just one hidden layer, given enough neurons, can approximate any continuous function to arbitrary precision. So in theory, one hidden layer suffices.
In practice it doesn't, for two reasons:
1. "Enough neurons" might mean an astronomical number — wider isn't free.
2. You'd need an astronomical amount of data, perfect optimisation, and infinite patience to actually find those weights.
Depth turns out to be cheaper than width. A 10-layer network with modest width can learn things a 1-layer network of any practical width cannot. This empirical fact — that depth beats width given a fixed parameter budget — is the entire reason the field is called deep learning.
Common Mistakes
- Thinking neural networks reason like humans. They don't. They are statistical pattern matchers. They have no concept of cause, agency, or "understanding". A model that scores 99% on a benchmark can fail on a slightly out-of-distribution example a child would handle.
- Reaching for NNs on small tabular data. Gradient-boosted trees (XGBoost, LightGBM) routinely beat neural networks on structured data with under a million rows. NNs shine on unstructured data — images, audio, text — and on huge datasets. See ml-pipeline for when to pick which.
- Ignoring the activation function. A network with no activations (or only linear ones) is mathematically equivalent to a single linear regression. Non-linearity is not optional.
- Forgetting that initial weights are random. Two training runs with different seeds produce slightly different models. This is normal — set a seed if you need reproducibility.
🎯 Your Turn — Compute a Neuron by Hand
Given this single-neuron model:
y = ReLU(0.5·x₁ + 0.3·x₂ − 0.2)
Compute the output y when x₁ = 2 and x₂ = 4.
Skeleton:
# Step 1: compute the weighted sum weighted_sum = ___ # Step 2: add the bias (already inside the formula above) # Step 3: apply ReLU y = ___
setup added so this can run · defines ___
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) ___ = _AutoMock('___')
Hint 1 — Substitute carefully
Multiply each input by its weight, then add the bias of −0.2. Don't skip the sign of the bias.Hint 2 — What ReLU does
ReLU returns the input if it's positive, otherwise zero. SoReLU(2.4) = 2.4 and ReLU(−0.5) = 0.
Show full solution
weighted_sum = 0.5 · 2 + 0.3 · 4 − 0.2 = 1.0 + 1.2 − 0.2 = 2.0 y = ReLU(2.0) = 2.0
That's it. One neuron, one forward pass, by hand. Stack a few thousand of these into layers, run a few million of these forward passes per second on a GPU, and you have a neural network.
Notice how boring the arithmetic is — that's the point. The intelligence isn't in any individual operation; it's in the pattern of weights that the training loop discovers.
What You Learned
- A neuron computes weighted sum → activation. Activations add the non-linearity that makes deep learning possible.
- ReLU for hidden layers, sigmoid for binary output, softmax for multi-class output.
- A layer is a group of neurons sharing inputs. A network is a stack of layers — deeper layers learn more abstract features.
- Training is a loop: forward pass → loss → backprop → weight update, repeat. Gradient descent is rolling a ball down the loss landscape.
- Neural networks are universal approximators in theory, but in practice depth, data, and compute are what make them work.
Next: Building Neural Networks with Keras — turning the concepts above into 10 lines of working code.