PythonMastery
beginner 22 min read · lesson 1 of 7 in Deep Learning Fundamentals

How Neural Networks Think

1 · The lesson

read

A neural network is a function. Numbers go in, numbers come out, and the function adjusts its own internal parameters until the outputs look like what you wanted. That sentence covers ninety-five percent of what's really happening inside the largest model running on a billion-dollar GPU cluster.

Everything else — convolutions, attention, transformers, diffusion — is a shape of that same machinery. So before any framework code, you need a clear mental model of the smallest unit (a neuron), how they stack into layers, how they learn, and where the metaphor with biology politely stops being useful.

This lesson is concept-first. No Keras yet. By the end you should be able to compute a neuron's output on paper.


1. The Smallest Unit — One Neuron

A neuron takes some numbers, multiplies each by a weight, adds a bias, and passes the result through an activation function.

python
y = activation(w₁·x₁ + w₂·x₂ + ... + wₙ·xₙ + b)

That's the whole thing. Two inputs:

python
y = activation(w₁·x₁ + w₂·x₂ + b)

The weights say how much each input matters. The bias shifts the result up or down so the neuron can fire even when all inputs are zero. The activation function adds the non-linearity that lets a stack of these units learn curves, not just straight lines.

Without an activation function, no matter how many neurons you stack, the whole network collapses into one big linear equation. The activation is what makes deep learning deep.


2. Activation Functions — The Four You'll See

Every framework ships dozens; you'll use four 95% of the time.

ActivationFormula (intuition)When to use
ReLUmax(0, x) — pass positives, kill negativesDefault for hidden layers
SigmoidSquashes to (0, 1) — an S-curveBinary classification output
TanhSquashes to (-1, 1) — centred S-curveRNN hidden states (historical)
SoftmaxTurns a vector into probabilities that sum to 1Multi-class classification output

ReLU is the modern workhorse. It's almost embarrassingly simple — return the input if positive, otherwise zero — and yet it solved the vanishing gradient problem that crippled sigmoid-based networks in the 1990s. Use ReLU for hidden layers unless you have a specific reason not to.

Sigmoid was historically used everywhere. Now it lives mostly in the final layer of binary classifiers (output ≈ probability the answer is "yes").

Tanh is sigmoid's cousin, but centred at zero. It mostly hangs around in older RNN architectures.

Softmax is special: it operates on a vector of numbers, not a single scalar. It exponentiates each value and divides by the sum, producing a probability distribution. If your model outputs 10 classes, softmax converts the 10 raw scores ("logits") into 10 probabilities summing to 1.


3. Stacking Neurons — Layers

A layer is a group of neurons that all see the same inputs. Each neuron in the layer has its own weights and bias, so each computes a slightly different function of those inputs.

A typical feed-forward network has three kinds of layers:

  • Input layer — just a placeholder for your features. No neurons, no learning.
  • Hidden layers — the real workers. One, two, twenty, a hundred. Depth is the number of these.
  • Output layer — produces the final prediction. Its size and activation depend on the task (1 + sigmoid for binary, N + softmax for N-class, 1 + linear for regression).

Each hidden layer takes the previous layer's outputs as its inputs. Numbers flow forward, layer by layer, getting transformed at every step.


4. Why Depth Helps — Hierarchical Features

Here is the intuition that justifies deep learning:

  • Layer 1 learns very simple patterns from raw input — for images, things like edges and colour gradients.
  • Layer 2 combines those simple patterns into slightly more complex ones — textures, corners, simple shapes.
  • Layer 3 combines those into still-more-complex parts — wheels, eyes, fur.
  • Final layers combine parts into whole-object detectors — "car", "cat", "stop sign".

You didn't tell the network what an edge or a wheel is. It discovered those intermediate concepts because they were useful for reducing its error. Each layer abstracts the layer below it. This is what people mean when they say neural networks "learn representations".

The same hierarchy shows up in language models (characters → words → phrases → meaning), audio models (waveforms → phonemes → words), and protein models (atoms → amino acids → folds).

The practical consequence: with images, audio, and text — domains where raw inputs are far from the labels — depth helps enormously. On clean tabular data, where each column is already a meaningful feature, a shallow model or a tree ensemble often wins. Depth pays off when the model needs to invent its own intermediate concepts.


5. The Forward Pass

You feed inputs into the input layer. Each hidden layer computes its outputs from the previous layer's outputs. Eventually the output layer produces a prediction.

python
input  →  hidden 1  →  hidden 2  →  output
 x         h₁          h₂            ŷ

That's the forward pass. It's pure arithmetic — multiply, add, activate, multiply, add, activate, all the way through. No learning happens here. You're just using the current weights to produce a prediction.

Initially the weights are random, so the prediction will be nonsense. A freshly-initialised binary classifier outputs roughly 0.5 for everything. A freshly-initialised 10-class softmax outputs roughly 0.1 for every class. The learning comes next.


6. The Loss Function — Measuring Wrongness

The loss is a number that says how wrong the prediction was compared to the truth. Lower is better. The whole training process is one giant exercise in pushing the loss downward.

You pick a loss function based on your task:

TaskLossWhy
Regression (predict a number)Mean Squared Error — average of (ŷ − y)²Penalises big errors quadratically
Binary classificationBinary cross-entropyMeasures probability distance
Multi-class classificationCategorical cross-entropyProbability distance over N classes

For a single prediction with true label y and predicted probability ŷ, binary cross-entropy is:

python
loss = -(y · log(ŷ) + (1 − y) · log(1 − ŷ))

Don't memorise it. Know that it grows large when the model is confidently wrong (predicts 0.01 when the answer was 1) and shrinks toward zero when the model is confidently right.


7. Backpropagation in Plain English

Once you have the loss, you need to know how to change the weights to reduce it. That's what backpropagation computes.

Imagine the network as a long chain of operations. Backprop walks that chain backwards from the loss, asking at each step: "if I nudged this weight a tiny bit, how much would the loss change?" The answer is the gradient for that weight.

You don't compute backprop by hand — every framework does it automatically (this is called autograd). But you should know what it produces: one gradient per weight, telling you the direction and magnitude to nudge that weight to lower the loss.


8. Gradient Descent — Rolling the Ball Downhill

Picture the loss as a bumpy landscape. Each weight is one dimension; a network with a million weights lives in a million-dimensional landscape. You can't visualise it, but the math doesn't care.

You're standing at a random point (the random initial weights). The gradient tells you which direction is steepest downhill. You take a small step in that direction:

python
new_weight = old_weight − learning_rate · gradient

The learning rate is the size of your step. Too big, you overshoot the valley. Too small, you take forever to get anywhere. We'll cover tuning it in Training & Improving.

Do this for every weight, with every batch of training data, for many passes through the dataset, and the weights gradually settle into a configuration where the loss is low.


9. The Training Loop in One Paragraph

For each batch of data: run the forward pass to get predictions, compute the loss against the true labels, run backpropagation to get the gradient of the loss with respect to every weight, then update every weight a small step in the downhill direction. Repeat until the loss stops dropping. That single paragraph is the engine of every neural network ever trained.

Some vocabulary that pins this down:

  • A batch is a small group of training examples (commonly 32 or 64) processed together.
  • One full pass through the training set is an epoch. You usually train for tens to hundreds of epochs.
  • The number of weight updates is (training_set_size / batch_size) · epochs. For a million-example dataset, batch 32, 10 epochs: 312,500 updates.

Each update is tiny. Convergence comes from doing huge numbers of them.


10. The Universal Function Approximator

A celebrated 1989 theorem says a neural network with just one hidden layer, given enough neurons, can approximate any continuous function to arbitrary precision. So in theory, one hidden layer suffices.

In practice it doesn't, for two reasons:

1. "Enough neurons" might mean an astronomical number — wider isn't free.
2. You'd need an astronomical amount of data, perfect optimisation, and infinite patience to actually find those weights.

Depth turns out to be cheaper than width. A 10-layer network with modest width can learn things a 1-layer network of any practical width cannot. This empirical fact — that depth beats width given a fixed parameter budget — is the entire reason the field is called deep learning.


Common Mistakes

  • Thinking neural networks reason like humans. They don't. They are statistical pattern matchers. They have no concept of cause, agency, or "understanding". A model that scores 99% on a benchmark can fail on a slightly out-of-distribution example a child would handle.
  • Reaching for NNs on small tabular data. Gradient-boosted trees (XGBoost, LightGBM) routinely beat neural networks on structured data with under a million rows. NNs shine on unstructured data — images, audio, text — and on huge datasets. See ml-pipeline for when to pick which.
  • Ignoring the activation function. A network with no activations (or only linear ones) is mathematically equivalent to a single linear regression. Non-linearity is not optional.
  • Forgetting that initial weights are random. Two training runs with different seeds produce slightly different models. This is normal — set a seed if you need reproducibility.

🎯 Your Turn — Compute a Neuron by Hand

Given this single-neuron model:

python
y = ReLU(0.5·x₁ + 0.3·x₂ − 0.2)

Compute the output y when x₁ = 2 and x₂ = 4.

Skeleton:

python
# Step 1: compute the weighted sum
weighted_sum = ___

# Step 2: add the bias (already inside the formula above)
# Step 3: apply ReLU
y = ___
+ setup added so this can run · defines ___
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

___ = _AutoMock('___')
Hint 1 — Substitute carefully Multiply each input by its weight, then add the bias of −0.2. Don't skip the sign of the bias.
Hint 2 — What ReLU does ReLU returns the input if it's positive, otherwise zero. So ReLU(2.4) = 2.4 and ReLU(−0.5) = 0.
Show full solution
python
weighted_sum = 0.5 · 2  +  0.3 · 4  −  0.2
             = 1.0      +  1.2      −  0.2
             = 2.0

y = ReLU(2.0) = 2.0

That's it. One neuron, one forward pass, by hand. Stack a few thousand of these into layers, run a few million of these forward passes per second on a GPU, and you have a neural network.

Notice how boring the arithmetic is — that's the point. The intelligence isn't in any individual operation; it's in the pattern of weights that the training loop discovers.


What You Learned

  • A neuron computes weighted sum → activation. Activations add the non-linearity that makes deep learning possible.
  • ReLU for hidden layers, sigmoid for binary output, softmax for multi-class output.
  • A layer is a group of neurons sharing inputs. A network is a stack of layers — deeper layers learn more abstract features.
  • Training is a loop: forward pass → loss → backprop → weight update, repeat. Gradient descent is rolling a ball down the loss landscape.
  • Neural networks are universal approximators in theory, but in practice depth, data, and compute are what make them work.

Next: Building Neural Networks with Keras — turning the concepts above into 10 lines of working code.