PythonMastery
beginner 20 min read · lesson 3 of 7 in Deep Learning Fundamentals

Building Neural Networks with Keras

1 · The lesson

read

Keras needs TensorFlow, which has no WebAssembly build, so the blocks below do not run in this tab — pip install tensorflow on your own machine. If you have not yet written the backward pass by hand, do A Network from Scratch first: every layer you are about to declare in one line, you will have already built in ten.

Heads up. TensorFlow does not run in this in-browser sandbox. Run these snippets in a Colab notebook (free GPU at colab.research.google.com) or locally after pip install tensorflow. Expected outputs are shown inline as comments.

Keras is the friendliest way to build a neural network. It sits on top of TensorFlow and turns the forward-pass / loss / backprop / update loop you met in the previous lesson into a handful of declarative lines. You describe the shape of the network; Keras handles every gradient.

If you've used scikit-learn's .fit() / .predict() rhythm, you already half-know Keras. Same vocabulary, deeper machinery.


1. Why Keras

Three things make Keras the canonical first DL framework:

  • Two-line model definition. A working multi-layer network fits on a postcard.
  • Sane defaults. Adam optimiser, He initialisation, sensible batch size — they all work out of the box.
  • Escape hatches. When you outgrow Sequential, the Functional API handles branches and multi-input/multi-output graphs. When you outgrow that, you drop down to raw TensorFlow.

PyTorch is the other dominant framework, more verbose but more flexible. We'll touch on it in the specialist track. For now, Keras.


2. The Two APIs

APIUse when
SequentialYou have a straight stack of layers — one input, one output, no branches
FunctionalYou need multiple inputs, multiple outputs, skip connections, or shared layers

Ninety percent of your first models will be Sequential. We'll start there.


3. A Minimal Sequential Model

python
from tensorflow import keras
from tensorflow.keras.layers import Dense
from tensorflow.keras.models import Sequential

model = Sequential([
    Dense(64, activation='relu', input_shape=(10,)),
    Dense(32, activation='relu'),
    Dense(1,  activation='sigmoid'),
])

model.compile(optimizer='adam',
              loss='binary_crossentropy',
              metrics=['accuracy'])

model.fit(X_train, y_train,
          epochs=10,
          batch_size=32,
          validation_split=0.2)
# → Epoch 10/10  loss: 0.23 - accuracy: 0.91 - val_loss: 0.27 - val_accuracy: 0.89
+ setup added so this can run · defines X_train, y_train
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_train = _AutoMock('X_train')
y_train = _AutoMock('y_train')

That's a complete binary classifier on 10-feature inputs. Three lines for the architecture, one to compile, one to train. Let's read it line by line.

  • Dense(64, activation='relu', input_shape=(10,)) — a fully connected layer with 64 neurons, ReLU activation, expecting inputs of length 10. The input_shape only needs to be specified on the first layer — Keras infers it everywhere else.
  • Dense(32, activation='relu') — a smaller hidden layer. Funnel-shaped networks (wider → narrower) are a common default.
  • Dense(1, activation='sigmoid') — one output neuron with sigmoid, producing a probability in (0, 1) for binary classification.

4. Compile — Three Choices

python
model.compile(optimizer='adam',
              loss='binary_crossentropy',
              metrics=['accuracy'])
+ setup added so this can run · defines model
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

model = _AutoMock('model')
  • optimizer — how weights get updated. 'adam' is the default for almost everything. You'll meet alternatives (SGD, RMSprop, AdamW) in the next lesson.
  • loss — what to minimise. Must match the task (see table in Section 8).
  • metrics — what to report during training, not what to optimise. Common choices: 'accuracy', 'mae', 'auc'. Pass a list — you can track several.

5. Fit — Train the Thing

python
history = model.fit(
    X_train, y_train,
    epochs=10,
    batch_size=32,
    validation_split=0.2,
    callbacks=[],
    verbose=1,
)
+ setup added so this can run · defines X_train, y_train, model
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_train = _AutoMock('X_train')
y_train = _AutoMock('y_train')
model = _AutoMock('model')
  • epochs — number of full passes through the training data. Start with 10–50 and use early stopping.
  • batch_size — how many samples to process before one weight update. 32 is the classic default. 64 / 128 / 256 are common for larger datasets.
  • validation_split=0.2 — Keras holds back 20% of X_train as a validation set. Use this or pass validation_data=(X_val, y_val) explicitly.
  • callbacks — list of objects that hook into training (early stopping, model checkpointing, learning rate scheduling). Covered in Training & Improving.
  • verbose — 1 for the progress bar, 2 for one line per epoch, 0 for silence.

fit() returns a History object whose .history dict holds per-epoch loss and metric values for both train and validation. You'll plot those obsessively.


6. Predict and Evaluate

python
# Evaluate on a held-out test set
test_loss, test_acc = model.evaluate(X_test, y_test, verbose=0)
print(f"test accuracy: {test_acc:.3f}")
# → test accuracy: 0.887

# Predict probabilities for new samples
probs = model.predict(X_new)
# → array shape (n_samples, 1), each entry in [0, 1]

# Turn probabilities into class labels (binary threshold)
preds = (probs > 0.5).astype(int)
+ setup added so this can run · defines X_test, y_test, X_new, model
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_test = _AutoMock('X_test')
y_test = _AutoMock('y_test')
X_new = _AutoMock('X_new')
model = _AutoMock('model')

predict() returns raw model output. For multi-class softmax models, use np.argmax(probs, axis=1) to get class indices.


7. Read the Summary

python
model.summary()
+ setup added so this can run · defines model
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

model = _AutoMock('model')
python
Model: "sequential"
_________________________________________________________________
 Layer (type)                Output Shape              Param #
=================================================================
 dense (Dense)               (None, 64)                704
 dense_1 (Dense)             (None, 32)                2080
 dense_2 (Dense)             (None, 1)                 33
=================================================================
Total params: 2,817
Trainable params: 2,817
Non-trainable params: 0
_________________________________________________________________

Two things to read here:

  • Output Shape — None is the batch dimension (variable). The number after it is the neuron count for that layer.
  • Param # — number of learnable weights. For a Dense layer it's (inputs · neurons) + neurons (the + neurons is the bias). First layer: 10·64 + 64 = 704. ✓

If a layer's param count surprises you, your input shape is probably wrong.


8. Output Layer & Loss — The Lookup Table

The single most common bug in a first Keras model is mismatching the final activation with the loss. Memorise this:

TaskFinal layerLoss
Binary classificationDense(1, activation='sigmoid')'binary_crossentropy'
Multi-class (single label)Dense(N, activation='softmax')'categorical_crossentropy' (one-hot labels) or 'sparse_categorical_crossentropy' (integer labels)
Multi-label classificationDense(N, activation='sigmoid')'binary_crossentropy'
RegressionDense(1, activation='linear') (or omit activation)'mse' or 'mae'

Use sparse_categorical_crossentropy when your labels are integers like [0, 3, 1, 2, ...]. Use plain categorical_crossentropy when they're one-hot encoded like [[1,0,0,0], [0,0,0,1], ...]. Same math, different label format.


9. Choosing Layer Sizes

There is no formula. Heuristics that won't waste your time:

  • Start small. Two hidden layers, sizes like (64, 32), are enough to beat a logistic regression on most datasets.
  • Funnel narrows. Each hidden layer is usually the same size as or smaller than the one before.
  • Powers of two. 32, 64, 128, 256 — GPUs are happy with these. Not load-bearing, just convention.
  • More data → wider/deeper is safer. Tiny dataset + huge network = guaranteed overfit.

10. The Functional API — One Example

Sequential can't handle two inputs merging into one output. Functional can:

python
from tensorflow.keras import Input, Model
from tensorflow.keras.layers import Dense, concatenate

# Two separate input branches
numeric_input = Input(shape=(10,), name='numeric')
category_input = Input(shape=(5,),  name='category')

n = Dense(32, activation='relu')(numeric_input)
c = Dense(16, activation='relu')(category_input)

merged = concatenate([n, c])
hidden = Dense(32, activation='relu')(merged)
output = Dense(1, activation='sigmoid')(hidden)

model = Model(inputs=[numeric_input, category_input], outputs=output)
model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])

model.fit({'numeric': X_num, 'category': X_cat}, y, epochs=10)
+ setup added so this can run · defines y, X_num, X_cat
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

y = _AutoMock('y')
X_num = _AutoMock('X_num')
X_cat = _AutoMock('X_cat')

Notice the pattern: every layer is called on the previous tensor, like Dense(32)(input). You assemble a graph by hand, then wrap it in Model(inputs=..., outputs=...). Compile and fit work identically.

You won't need this for the next three lessons, but it's worth knowing it exists.


11. Save and Load

python
# Save — single file containing architecture + weights + optimiser state
model.save("classifier.keras")

# Load — anywhere, including a different machine
restored = keras.models.load_model("classifier.keras")
restored.predict(X_new)
+ setup added so this can run · defines X_new, model, keras
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_new = _AutoMock('X_new')
model = _AutoMock('model')
keras = _AutoMock('keras')

The .keras extension is the modern unified format. Old code uses .h5 (Keras HDF5) or SavedModel/ directories — both still work, but .keras is the recommended default since 2023.

We cover production serving in Deploy Your Neural Network.


Common Mistakes

  • Wrong loss for the task. Using 'mse' on a classification problem will train, but it's a sign you've copy-pasted from the wrong tutorial. Match the Section 8 table.
  • Wrong final activation. Multi-class without softmax → outputs aren't probabilities → cross-entropy gives garbage. Linear in a binary classifier → predictions can be anywhere on the number line.
  • Forgetting input_shape on the first layer. Without it, the model isn't built until you call fit, so model.summary() errors out with "this model has not yet been built".
  • Network too small. 4 neurons trying to learn MNIST will plateau at chance. If your loss won't drop, double the width before blaming the data.
  • Network too big for the data. 10,000 neurons on 200 training samples is a memorisation machine, not a learner.

🎯 Your Turn — Build a Binary Classifier

You have a dataset with 20 features and a binary label. Build a three-layer Sequential model — hidden layers of 64 then 32 neurons (ReLU), output layer for binary classification. Compile with Adam, binary cross-entropy, and accuracy. Print the model summary.

Skeleton:

python
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense

model = Sequential([
    # TODO 1: first hidden layer — 64 neurons, ReLU, input_shape for 20 features
    # TODO 2: second hidden layer — 32 neurons, ReLU
    # TODO 3: output layer — binary classification
])

# TODO 4: compile with Adam, binary cross-entropy, accuracy

model.summary()
Hint 1 — input_shape on layer 1 The first Dense layer needs input_shape=(20,). The trailing comma is required — it's a one-element tuple, not a number in parentheses.
Hint 2 — Binary output One output neuron with activation='sigmoid', and loss='binary_crossentropy' at compile time. Don't reach for softmax on binary — sigmoid is correct.
Show full solution
python
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import Dense

model = Sequential([
    Dense(64, activation='relu', input_shape=(20,)),
    Dense(32, activation='relu'),
    Dense(1,  activation='sigmoid'),
])

model.compile(optimizer='adam',
              loss='binary_crossentropy',
              metrics=['accuracy'])

model.summary()
# → Total params: 3,489
#   layer 1: 20·64 + 64  = 1,344
#   layer 2: 64·32 + 32  = 2,080
#   layer 3: 32·1  + 1   =    33

To train it, you'd add:

python
history = model.fit(X_train, y_train,
                    epochs=20,
                    batch_size=32,
                    validation_split=0.2)
# → Epoch 20/20  loss: 0.31 - accuracy: 0.88 - val_loss: 0.34 - val_accuracy: 0.86
+ setup added so this can run · defines X_train, y_train, model
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_train = _AutoMock('X_train')
y_train = _AutoMock('y_train')
model = _AutoMock('model')

What You Learned

  • Sequential for straight stacks; Functional when you need branches.
  • A Keras model is define → compile → fit → evaluate → predict.
  • The final activation and the loss must match the task — keep the Section 8 table close.
  • model.summary() is your first sanity check; the Param # column tells you whether the shapes are right.
  • Save with model.save("name.keras"), load with keras.models.load_model("name.keras").

Next: Training & Improving Your Network — diagnosing what's going wrong and the four tools you'll reach for to fix it.