Building Neural Networks with Keras
1 · The lesson
readKeras needs TensorFlow, which has no WebAssembly build, so the blocks below do not run in this tab —
pip install tensorflowon your own machine. If you have not yet written the backward pass by hand, do A Network from Scratch first: every layer you are about to declare in one line, you will have already built in ten.
Heads up. TensorFlow does not run in this in-browser sandbox. Run these snippets in a Colab notebook (free GPU at
colab.research.google.com) or locally afterpip install tensorflow. Expected outputs are shown inline as comments.
Keras is the friendliest way to build a neural network. It sits on top of TensorFlow and turns the forward-pass / loss / backprop / update loop you met in the previous lesson into a handful of declarative lines. You describe the shape of the network; Keras handles every gradient.
If you've used scikit-learn's .fit() / .predict() rhythm, you already half-know Keras. Same vocabulary, deeper machinery.
1. Why Keras
Three things make Keras the canonical first DL framework:
- Two-line model definition. A working multi-layer network fits on a postcard.
- Sane defaults. Adam optimiser, He initialisation, sensible batch size — they all work out of the box.
- Escape hatches. When you outgrow
Sequential, the Functional API handles branches and multi-input/multi-output graphs. When you outgrow that, you drop down to raw TensorFlow.
PyTorch is the other dominant framework, more verbose but more flexible. We'll touch on it in the specialist track. For now, Keras.
2. The Two APIs
| API | Use when |
|---|---|
| Sequential | You have a straight stack of layers — one input, one output, no branches |
| Functional | You need multiple inputs, multiple outputs, skip connections, or shared layers |
Ninety percent of your first models will be Sequential. We'll start there.
3. A Minimal Sequential Model
from tensorflow import keras from tensorflow.keras.layers import Dense from tensorflow.keras.models import Sequential model = Sequential([ Dense(64, activation='relu', input_shape=(10,)), Dense(32, activation='relu'), Dense(1, activation='sigmoid'), ]) model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy']) model.fit(X_train, y_train, epochs=10, batch_size=32, validation_split=0.2) # → Epoch 10/10 loss: 0.23 - accuracy: 0.91 - val_loss: 0.27 - val_accuracy: 0.89
setup added so this can run · defines X_train, y_train
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_train = _AutoMock('X_train') y_train = _AutoMock('y_train')
That's a complete binary classifier on 10-feature inputs. Three lines for the architecture, one to compile, one to train. Let's read it line by line.
Dense(64, activation='relu', input_shape=(10,))— a fully connected layer with 64 neurons, ReLU activation, expecting inputs of length 10. Theinput_shapeonly needs to be specified on the first layer — Keras infers it everywhere else.Dense(32, activation='relu')— a smaller hidden layer. Funnel-shaped networks (wider → narrower) are a common default.Dense(1, activation='sigmoid')— one output neuron with sigmoid, producing a probability in(0, 1)for binary classification.
4. Compile — Three Choices
model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy'])
setup added so this can run · defines model
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) model = _AutoMock('model')
optimizer— how weights get updated.'adam'is the default for almost everything. You'll meet alternatives (SGD, RMSprop, AdamW) in the next lesson.loss— what to minimise. Must match the task (see table in Section 8).metrics— what to report during training, not what to optimise. Common choices:'accuracy','mae','auc'. Pass a list — you can track several.
5. Fit — Train the Thing
history = model.fit( X_train, y_train, epochs=10, batch_size=32, validation_split=0.2, callbacks=[], verbose=1, )
setup added so this can run · defines X_train, y_train, model
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_train = _AutoMock('X_train') y_train = _AutoMock('y_train') model = _AutoMock('model')
epochs— number of full passes through the training data. Start with 10–50 and use early stopping.batch_size— how many samples to process before one weight update. 32 is the classic default. 64 / 128 / 256 are common for larger datasets.validation_split=0.2— Keras holds back 20% ofX_trainas a validation set. Use this or passvalidation_data=(X_val, y_val)explicitly.callbacks— list of objects that hook into training (early stopping, model checkpointing, learning rate scheduling). Covered in Training & Improving.verbose—1for the progress bar,2for one line per epoch,0for silence.
fit() returns a History object whose .history dict holds per-epoch loss and metric values for both train and validation. You'll plot those obsessively.
6. Predict and Evaluate
# Evaluate on a held-out test set test_loss, test_acc = model.evaluate(X_test, y_test, verbose=0) print(f"test accuracy: {test_acc:.3f}") # → test accuracy: 0.887 # Predict probabilities for new samples probs = model.predict(X_new) # → array shape (n_samples, 1), each entry in [0, 1] # Turn probabilities into class labels (binary threshold) preds = (probs > 0.5).astype(int)
setup added so this can run · defines X_test, y_test, X_new, model
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_test = _AutoMock('X_test') y_test = _AutoMock('y_test') X_new = _AutoMock('X_new') model = _AutoMock('model')
predict() returns raw model output. For multi-class softmax models, use np.argmax(probs, axis=1) to get class indices.
7. Read the Summary
model.summary() setup added so this can run · defines model
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) model = _AutoMock('model')
Model: "sequential" _________________________________________________________________ Layer (type) Output Shape Param # ================================================================= dense (Dense) (None, 64) 704 dense_1 (Dense) (None, 32) 2080 dense_2 (Dense) (None, 1) 33 ================================================================= Total params: 2,817 Trainable params: 2,817 Non-trainable params: 0 _________________________________________________________________
Two things to read here:
- Output Shape —
Noneis the batch dimension (variable). The number after it is the neuron count for that layer. - Param # — number of learnable weights. For a
Denselayer it's(inputs · neurons) + neurons(the+ neuronsis the bias). First layer:10·64 + 64 = 704. ✓
If a layer's param count surprises you, your input shape is probably wrong.
8. Output Layer & Loss — The Lookup Table
The single most common bug in a first Keras model is mismatching the final activation with the loss. Memorise this:
| Task | Final layer | Loss |
|---|---|---|
| Binary classification | Dense(1, activation='sigmoid') | 'binary_crossentropy' |
| Multi-class (single label) | Dense(N, activation='softmax') | 'categorical_crossentropy' (one-hot labels) or 'sparse_categorical_crossentropy' (integer labels) |
| Multi-label classification | Dense(N, activation='sigmoid') | 'binary_crossentropy' |
| Regression | Dense(1, activation='linear') (or omit activation) | 'mse' or 'mae' |
Use sparse_categorical_crossentropy when your labels are integers like [0, 3, 1, 2, ...]. Use plain categorical_crossentropy when they're one-hot encoded like [[1,0,0,0], [0,0,0,1], ...]. Same math, different label format.
9. Choosing Layer Sizes
There is no formula. Heuristics that won't waste your time:
- Start small. Two hidden layers, sizes like
(64, 32), are enough to beat a logistic regression on most datasets. - Funnel narrows. Each hidden layer is usually the same size as or smaller than the one before.
- Powers of two. 32, 64, 128, 256 — GPUs are happy with these. Not load-bearing, just convention.
- More data → wider/deeper is safer. Tiny dataset + huge network = guaranteed overfit.
10. The Functional API — One Example
Sequential can't handle two inputs merging into one output. Functional can:
from tensorflow.keras import Input, Model from tensorflow.keras.layers import Dense, concatenate # Two separate input branches numeric_input = Input(shape=(10,), name='numeric') category_input = Input(shape=(5,), name='category') n = Dense(32, activation='relu')(numeric_input) c = Dense(16, activation='relu')(category_input) merged = concatenate([n, c]) hidden = Dense(32, activation='relu')(merged) output = Dense(1, activation='sigmoid')(hidden) model = Model(inputs=[numeric_input, category_input], outputs=output) model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy']) model.fit({'numeric': X_num, 'category': X_cat}, y, epochs=10)
setup added so this can run · defines y, X_num, X_cat
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) y = _AutoMock('y') X_num = _AutoMock('X_num') X_cat = _AutoMock('X_cat')
Notice the pattern: every layer is called on the previous tensor, like Dense(32)(input). You assemble a graph by hand, then wrap it in Model(inputs=..., outputs=...). Compile and fit work identically.
You won't need this for the next three lessons, but it's worth knowing it exists.
11. Save and Load
# Save — single file containing architecture + weights + optimiser state model.save("classifier.keras") # Load — anywhere, including a different machine restored = keras.models.load_model("classifier.keras") restored.predict(X_new)
setup added so this can run · defines X_new, model, keras
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_new = _AutoMock('X_new') model = _AutoMock('model') keras = _AutoMock('keras')
The .keras extension is the modern unified format. Old code uses .h5 (Keras HDF5) or SavedModel/ directories — both still work, but .keras is the recommended default since 2023.
We cover production serving in Deploy Your Neural Network.
Common Mistakes
- Wrong loss for the task. Using
'mse'on a classification problem will train, but it's a sign you've copy-pasted from the wrong tutorial. Match the Section 8 table. - Wrong final activation. Multi-class without softmax → outputs aren't probabilities → cross-entropy gives garbage. Linear in a binary classifier → predictions can be anywhere on the number line.
- Forgetting
input_shapeon the first layer. Without it, the model isn't built until you callfit, somodel.summary()errors out with "this model has not yet been built". - Network too small. 4 neurons trying to learn MNIST will plateau at chance. If your loss won't drop, double the width before blaming the data.
- Network too big for the data. 10,000 neurons on 200 training samples is a memorisation machine, not a learner.
🎯 Your Turn — Build a Binary Classifier
You have a dataset with 20 features and a binary label. Build a three-layer Sequential model — hidden layers of 64 then 32 neurons (ReLU), output layer for binary classification. Compile with Adam, binary cross-entropy, and accuracy. Print the model summary.
Skeleton:
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense model = Sequential([ # TODO 1: first hidden layer — 64 neurons, ReLU, input_shape for 20 features # TODO 2: second hidden layer — 32 neurons, ReLU # TODO 3: output layer — binary classification ]) # TODO 4: compile with Adam, binary cross-entropy, accuracy model.summary()
Hint 1 — input_shape on layer 1
The firstDense layer needs input_shape=(20,). The trailing comma is required — it's a one-element tuple, not a number in parentheses.
Hint 2 — Binary output
One output neuron withactivation='sigmoid', and loss='binary_crossentropy' at compile time. Don't reach for softmax on binary — sigmoid is correct.
Show full solution
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense model = Sequential([ Dense(64, activation='relu', input_shape=(20,)), Dense(32, activation='relu'), Dense(1, activation='sigmoid'), ]) model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy']) model.summary() # → Total params: 3,489 # layer 1: 20·64 + 64 = 1,344 # layer 2: 64·32 + 32 = 2,080 # layer 3: 32·1 + 1 = 33
To train it, you'd add:
history = model.fit(X_train, y_train, epochs=20, batch_size=32, validation_split=0.2) # → Epoch 20/20 loss: 0.31 - accuracy: 0.88 - val_loss: 0.34 - val_accuracy: 0.86
setup added so this can run · defines X_train, y_train, model
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_train = _AutoMock('X_train') y_train = _AutoMock('y_train') model = _AutoMock('model')
What You Learned
- Sequential for straight stacks; Functional when you need branches.
- A Keras model is define → compile → fit → evaluate → predict.
- The final activation and the loss must match the task — keep the Section 8 table close.
model.summary()is your first sanity check; theParam #column tells you whether the shapes are right.- Save with
model.save("name.keras"), load withkeras.models.load_model("name.keras").
Next: Training & Improving Your Network — diagnosing what's going wrong and the four tools you'll reach for to fix it.