PythonMastery
advanced 26 min read · lesson 4 of 9 in AI & Deep Learning

RNNs & LSTMs: Sequential Architectures

1 · The lesson

read

A CNN looks at all of an image at once. A transformer looks at all of a sequence at once. An RNN looks at one token at a time, carrying a hidden state forward through time — the only mainstream architecture that operates step by step on a sequence.

Transformers have replaced RNNs for almost all NLP since 2018. So why is this lesson here? Because RNNs are still the right call for streaming inference (token-by-token generation without re-processing the whole history), time series forecasting (small datasets, structured temporal correlations), and edge devices (tiny models, tiny memory). And because if you understand the LSTM, you understand half of what attention later replaced — which makes the transformer's design choices much clearer.


1. The Problem RNNs Solve

A standard MLP has a fixed input shape — say, a 224×224 image. But what's the "input shape" of a sentence? Sentences vary in length. Audio clips vary in length. Time series vary in length. You need an architecture that:

  • Accepts variable-length input,
  • Maintains state across time steps,
  • Shares parameters so the model doesn't grow with the sequence length.

RNNs do all three. The same set of weights is applied at every time step, with a hidden state passing information forward:

$$
\mathbf{h}_t = f(\mathbf{x}_t, \mathbf{h}_{t-1})
$$

where $\mathbf{x}_t$ is the input at time $t$ and $\mathbf{h}_t$ is the hidden state. The function $f$ has learned weights, but those weights are shared across all $t$. A 100-step sequence and a 10-step sequence use the same parameters — the RNN just unrolls $f$ for as many steps as you give it.


2. The Vanilla RNN

The simplest version:

$$
\mathbf{h}_t = \tanh(W_h \mathbf{h}_{t-1} + W_x \mathbf{x}_t + \mathbf{b})
$$

$$
\mathbf{y}_t = W_y \mathbf{h}_t + \mathbf{b}_y
$$

Three weight matrices: $W_h$ (state-to-state), $W_x$ (input-to-state), $W_y$ (state-to-output). Same matrices used at every $t$. The hidden state is the network's memory.

In code, conceptually:

python
def rnn_step(x_t, h_prev, W_h, W_x, b):
    return np.tanh(W_h @ h_prev + W_x @ x_t + b)

h = np.zeros(hidden_size)
for x_t in sequence:
    h = rnn_step(x_t, h, W_h, W_x, b)
+ setup added so this can run · defines sequence, hidden_size, np
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

sequence = ["alpha", "beta", "gamma"]
hidden_size = _AutoMock('hidden_size')
np = _AutoMock('np')

Training works the same as any other network — backprop, except the gradient flows backward through time (BPTT, backpropagation through time): from the final output, back through every time step, accumulating gradients on the shared weights at every step.


3. Vanishing & Exploding Gradients

BPTT through $T$ time steps involves multiplying $T$ Jacobians of the recurrence together. If the Jacobian's spectral norm is consistently less than 1, the product collapses to zero — vanishing gradient. The model can't learn long-range dependencies because the gradient signal can't reach back far enough. If the spectral norm is consistently more than 1, the product explodes — exploding gradient, NaNs in the loss.

For a vanilla RNN with tanh, the Jacobian is bounded by the spectral norm of $W_h$ times the slope of tanh (≤ 1). With reasonable weights, vanishing wins — by step 20 or 30, gradient information from the start of the sequence is dust.

Gradient clipping is the cheap fix for explosion: $\mathbf{g} \leftarrow \mathbf{g} \cdot \min(1, c / \|\mathbf{g}\|)$. One line in Keras (clipnorm=1.0 on the optimiser). Doesn't fix vanishing — that needs a structural change. Enter LSTM.


4. LSTM — The Long Short-Term Memory Cell

Hochreiter & Schmidhuber introduced the LSTM in 1997 to solve vanishing gradients. The key idea: alongside the hidden state, maintain a cell state $\mathbf{c}_t$ that flows through time with only minor modifications, controlled by gates that decide what to forget, what to write, and what to read.

For each time step, given input $\mathbf{x}_t$, previous hidden $\mathbf{h}_{t-1}$, previous cell $\mathbf{c}_{t-1}$:

$$
\mathbf{f}_t = \sigma(W_f [\mathbf{h}_{t-1}, \mathbf{x}_t] + \mathbf{b}_f) \qquad \text{forget gate}
$$

$$
\mathbf{i}_t = \sigma(W_i [\mathbf{h}_{t-1}, \mathbf{x}_t] + \mathbf{b}_i) \qquad \text{input gate}
$$

$$
\tilde{\mathbf{c}}_t = \tanh(W_c [\mathbf{h}_{t-1}, \mathbf{x}_t] + \mathbf{b}_c) \qquad \text{candidate cell update}
$$

$$
\mathbf{c}_t = \mathbf{f}_t \odot \mathbf{c}_{t-1} + \mathbf{i}_t \odot \tilde{\mathbf{c}}_t \qquad \text{new cell state}
$$

$$
\mathbf{o}_t = \sigma(W_o [\mathbf{h}_{t-1}, \mathbf{x}_t] + \mathbf{b}_o) \qquad \text{output gate}
$$

$$
\mathbf{h}_t = \mathbf{o}_t \odot \tanh(\mathbf{c}_t) \qquad \text{new hidden state}
$$

$\sigma$ is the sigmoid (gate values in $[0, 1]$); $\odot$ is element-wise multiply. The cell state $\mathbf{c}_t$ is updated by an addition rather than a multiplication. Adding doesn't compound to zero or infinity through time — the gradient can flow back along the cell state highway with little decay. That's the whole reason LSTMs work.

Conceptually:

GateJob
Forget $\mathbf{f}_t$What of the old cell state should we erase?
Input $\mathbf{i}_t$How much of the new candidate should we write in?
Output $\mathbf{o}_t$What of the cell state should we expose as hidden state?

Three sigmoid gates and one tanh candidate, all from the same $[\mathbf{h}_{t-1}, \mathbf{x}_t]$ input. Roughly 4× the parameter count of a vanilla RNN with the same hidden size.


5. GRU — The Simplified Cousin

Cho et al. (2014) introduced the GRU as a stripped-down LSTM. Two gates instead of three; no separate cell state — just the hidden state.

$$
\mathbf{z}_t = \sigma(W_z [\mathbf{h}_{t-1}, \mathbf{x}_t]) \qquad \text{update gate}
$$

$$
\mathbf{r}_t = \sigma(W_r [\mathbf{h}_{t-1}, \mathbf{x}_t]) \qquad \text{reset gate}
$$

$$
\tilde{\mathbf{h}}_t = \tanh(W [\mathbf{r}_t \odot \mathbf{h}_{t-1}, \mathbf{x}_t])
$$

$$
\mathbf{h}_t = (1 - \mathbf{z}_t) \odot \mathbf{h}_{t-1} + \mathbf{z}_t \odot \tilde{\mathbf{h}}_t
$$

The update gate $\mathbf{z}_t$ interpolates between keeping the old hidden state and writing the candidate. ~25% fewer params than LSTM; often equivalent performance on small/medium tasks. Try both, pick the one that trains better on your problem. Modern empirical consensus: roughly equivalent on most tasks; LSTM slightly better on very long sequences, GRU slightly faster.


6. Bidirectional RNNs

A unidirectional RNN at step $t$ has only seen tokens $1$ through $t$. For sequence tagging (NER, POS), each token's classification benefits from seeing the whole sentence. A bidirectional RNN runs two RNNs — one forward, one backward — and concatenates their hidden states:

$$
\mathbf{h}_t = [\overrightarrow{\mathbf{h}_t} \,;\, \overleftarrow{\mathbf{h}_t}]
$$

In Keras:

python
from tensorflow.keras import layers

x = layers.Bidirectional(layers.LSTM(64, return_sequences=True))(embedded)
# Output channel dim is 64 * 2 = 128
+ setup added so this can run · defines embedded
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

embedded = _AutoMock('embedded')

Use bidirectional when:

  • The whole sequence is available at inference time (text classification, NER).
  • The future context is useful.

Don't use bidirectional for streaming (you don't have the future yet) or for autoregressive generation (you'd be cheating).


7. Stacked RNNs

A single RNN layer captures one level of abstraction. Stack multiple layers, feeding the lower layer's hidden states as input to the upper:

python
x = layers.LSTM(128, return_sequences=True)(x)    # MUST return sequences
x = layers.LSTM(64,  return_sequences=True)(x)
x = layers.LSTM(32,  return_sequences=False)(x)   # final layer returns just the last state
+ setup added so this can run · defines layers
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

layers = _AutoMock('layers')

return_sequences=True is mandatory on all but the final layer of a stack. Without it, the layer returns only the final hidden state, and the next layer has nothing to chew on. This is the single most common LSTM bug.

Deeper stacks (3-4 layers) help on large datasets with rich structure. On small datasets, one bidirectional layer + a dense head is usually enough.


8. Sequence-to-Sequence

When the input and output are both sequences (translation, summarisation, speech recognition), you have seq2seq. Encoder reads the input, produces a fixed-size context vector. Decoder reads the context vector and generates the output one token at a time.

python
Input:  "The cat sat" →  [encoder LSTM] → context vector
                                                  ↓
Output: "Le chat est"  ← [decoder LSTM, conditioned on context]

During training, the decoder is fed the correct previous token at each step ("teacher forcing") — much faster than letting it use its own (probably wrong) predictions, especially early in training. At inference, you have no ground truth, so the decoder uses its own previous prediction (greedy decoding, beam search, sampling, etc.).

The problem with a fixed-size context: a single vector has to encode the entire input sequence, regardless of length. For a 50-word sentence, that's an information bottleneck. The fix:


9. Attention — The Pre-Transformer Breakthrough

Bahdanau et al. (2015): instead of compressing the input to one vector, let the decoder attend to all encoder hidden states at each step, weighting them by relevance.

At each decoder step $t$, given the current decoder state $\mathbf{s}_t$ and encoder states $\mathbf{H} = [\mathbf{h}_1, \ldots, \mathbf{h}_T]$:

$$
e_{t,i} = \text{score}(\mathbf{s}_t, \mathbf{h}_i) \qquad \text{e.g. } \mathbf{s}_t^\top W \mathbf{h}_i
$$

$$
\alpha_{t,i} = \text{softmax}_i(e_{t,i}) \qquad \text{alignment weights, summing to 1}
$$

$$
\mathbf{c}_t = \sum_i \alpha_{t,i} \mathbf{h}_i \qquad \text{context vector for step } t
$$

The decoder now uses $\mathbf{c}_t$ (which depends on the decoder's current state) instead of a fixed encoded context. Each decoder step looks at the most relevant parts of the input — translation models learn alignments that look remarkably like human word alignments.

This is exactly the same attention that transformers later generalised: query (decoder state), keys (encoder states), values (encoder states), softmax-weighted sum. The transformer's contribution wasn't "invent attention" — it was "drop the RNN entirely, use only attention."


10. Why Transformers Replaced RNNs

For NLP (and increasingly everywhere), RNNs are obsolete because:

AspectRNNTransformer
Training parallelismSequential — must process step $t-1$ before step $t$Parallel — all positions computed at once
Long-range dependenciesLSTM mitigates but still degrades past ~200 stepsDirect attention, $O(1)$ path between any two positions
GPU utilisationPoor — small matmuls per stepExcellent — large batched matmuls
Pretrained ecosystemLimitedDominant (BERT, GPT, T5, Llama...)
Memory at inferenceConstant per stepLinear or quadratic per step (KV cache)

The training parallelism is the killer. A transformer processes a 1000-token sequence in one matrix multiply per layer; an RNN needs 1000 sequential ops. On modern hardware, the transformer is 100× faster to train, and pretraining at scale is what unlocks LLMs.

See Transformer Architecture for the deep dive on attention, multi-head, position embeddings, and the modern stack.


11. When RNNs Are Still Right

The pendulum hasn't swung entirely — RNNs remain the better choice for:

Time series forecasting

For univariate or low-dim multivariate time series with strong temporal correlation (sales, sensor data, weather), LSTMs with 50-200 hidden units often beat transformer-based forecasters. Transformers are data-hungry; a 10k-point time series doesn't have enough data to train one well. Specialised time series transformers (Informer, Autoformer, PatchTST, TFT) are catching up, but for "I have one signal and want next-month forecasts," LSTM is still a strong default.

Streaming / token-by-token inference

An RNN at inference processes one token at a time with O(1) memory — just the hidden state. A vanilla transformer re-attends to the entire history at every step (the KV cache helps, but it grows with sequence length). For very long generation on memory-constrained devices, RNNs win.

Edge devices

A 50k-parameter LSTM runs on a microcontroller. The smallest useful transformer is millions of parameters. Speech keyword detection on phones, gesture recognition on watches, anomaly detection on industrial sensors — RNN/LSTM territory.

Practical heuristic

WorkloadDefault architecture
Text classification, NLP tasksPretrained transformer
Time series forecasting (small/medium data)LSTM / GRU, or a specialised TS transformer
Real-time streaming inference (edge, mobile)LSTM / GRU
Anything you'd ship to a microcontrollerTiny LSTM, or a CNN
Big-data sequence modelling (LLM-scale)Transformer, full stop

12. Keras LSTM in Practice

The full API surface you'll actually use:

python
from tensorflow.keras import layers

# Many-to-one (sequence classification)
x = layers.LSTM(64, return_sequences=False)(x)        # output: (B, 64)

# Many-to-many (per-step output: NER, tagging)
x = layers.LSTM(64, return_sequences=True)(x)         # output: (B, T, 64)

# Bidirectional
x = layers.Bidirectional(layers.LSTM(64, return_sequences=True))(x)

# Stacked
x = layers.LSTM(128, return_sequences=True)(x)
x = layers.LSTM(64,  return_sequences=False)(x)

# With dropout (recurrent dropout — masks recurrent connections)
x = layers.LSTM(64, dropout=0.2, recurrent_dropout=0.2)(x)

# Return the full state (for seq2seq decoder initialisation)
out, hidden, cell = layers.LSTM(64, return_state=True)(x)

Gradient clipping is essentially mandatory for RNNs:

python
opt = keras.optimizers.Adam(learning_rate=1e-3, clipnorm=1.0)
+ setup added so this can run · defines keras
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

keras = _AutoMock('keras')

clipnorm=1.0 rescales the gradient vector to have at most norm 1. Without it, exploding gradients give NaN losses within a few epochs on hard sequences.


13. A Time Series Example

Predict the next value from the previous 50:

python
import numpy as np
import tensorflow as tf
from tensorflow.keras import layers, Model, Input

# Toy series — sine wave with noise
t = np.linspace(0, 100, 5000)
y = np.sin(t) + 0.1 * np.random.randn(len(t))

# Windowed: 50 past values → next value
WINDOW = 50
X = np.array([y[i:i+WINDOW]   for i in range(len(y) - WINDOW - 1)])
Y = np.array([y[i+WINDOW]     for i in range(len(y) - WINDOW - 1)])
X = X[..., None]                       # (N, 50, 1)  — channels axis required

inputs = Input(shape=(WINDOW, 1))
x = layers.LSTM(32, return_sequences=False)(inputs)
x = layers.Dense(1)(x)

model = Model(inputs, x)
model.compile(optimizer=tf.keras.optimizers.Adam(1e-3, clipnorm=1.0),
              loss="mse", metrics=["mae"])
model.fit(X, Y, epochs=10, batch_size=32, validation_split=0.2)

# Expected: MAE ~0.08-0.12 (close to the noise floor of 0.1)

Run in Colab or locally with pip install tensorflow. Expected outputs in comments.

This is the canonical recipe. The hidden-state size, window length, and depth are the dials you turn. For multi-step forecasting, switch to seq2seq (encoder reads the window, decoder outputs N future steps) or train one model per horizon.


14. Text Generation — The Char-Level LSTM

Karpathy's 2015 demo: train a char-level LSTM on Shakespeare, sample from it, get plausible-sounding output. Still a great teaching example.

python
# Pseudo-code — the structure, not all the plumbing
inputs = Input(shape=(SEQ_LEN,))
x = layers.Embedding(vocab_size, 128)(inputs)
x = layers.LSTM(256, return_sequences=True)(x)
x = layers.Dense(vocab_size)(x)           # logits per position
model = Model(inputs, x)
model.compile(
    optimizer=tf.keras.optimizers.Adam(1e-3, clipnorm=1.0),
    loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
)
+ setup added so this can run · defines Input, Model, vocab_size, SEQ_LEN, layers, tf
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def Input(*_a, **_kw):
    print('-> Input() called')
    return _AutoMock('Input()')
def Model(*_a, **_kw):
    print('-> Model() called')
    return _AutoMock('Model()')
vocab_size = _AutoMock('vocab_size')
SEQ_LEN = _AutoMock('SEQ_LEN')
layers = _AutoMock('layers')
tf = _AutoMock('tf')

Temperature sampling at generation time: divide logits by a temperature $\tau$ before softmax.

$$
P_i = \frac{\exp(z_i / \tau)}{\sum_j \exp(z_j / \tau)}
$$

  • $\tau < 1$: sharper distribution. Output is more confident, more repetitive, "safer".
  • $\tau = 1$: model's native distribution.
  • $\tau > 1$: flatter distribution. Output is more diverse, more creative, also more incoherent.
python
def sample(logits, temperature=1.0):
    logits = logits / temperature
    probs  = tf.nn.softmax(logits).numpy()
    return np.random.choice(len(probs), p=probs)
+ setup added so this can run · defines np, tf
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

np = _AutoMock('np')
tf = _AutoMock('tf')

Top-k and nucleus (top-p) sampling — variants that truncate the distribution before sampling — are the modern defaults in LLMs. The temperature dial is the same.


15. Masking — Variable-Length Sequences

Real sequences come in variable lengths. You batch them by padding to the same length, but the padding tokens shouldn't influence the loss or the hidden state. Masking tells the RNN which positions to ignore.

python
# Embedding layer with mask_zero=True propagates a mask for any zero-valued input
x = layers.Embedding(vocab_size, 128, mask_zero=True)(inputs)
x = layers.LSTM(64)(x)         # automatically respects the propagated mask
+ setup added so this can run · defines inputs, vocab_size, layers
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

inputs = _AutoMock('inputs')
vocab_size = _AutoMock('vocab_size')
layers = _AutoMock('layers')

When you pad with 0, the mask layer marks those positions and the LSTM skips them. Downstream Keras layers (LSTM, Bidirectional, Dense with sparse_categorical_crossentropy) all honour the mask. mask_zero=True on the Embedding is the one-line solution for variable-length sequences in Keras.

If you need explicit control, use layers.Masking(mask_value=0.0) before the LSTM.


Common Mistakes

1. Forgetting return_sequences=True on a stacked LSTM

python
x = layers.LSTM(128)(x)            # returns (B, 128) — just the final state
x = layers.LSTM(64)(x)             # error or silently wrong: expects 3D
+ setup added so this can run · defines layers
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

layers = _AutoMock('layers')

The second LSTM needs a 3D input (batch, time, features). Without return_sequences=True on the first, you've thrown away the time dimension. Bug counter goes up by one every time someone learns this.

2. Comparing LSTM val loss without seeding

LSTMs have high run-to-run variance. The difference between two model variants can be smaller than the noise across random seeds. Set tf.keras.utils.set_random_seed(42) and report mean ± std over 3-5 seeds, or you're measuring noise.

3. Using LSTM for sequences past a few hundred tokens

LSTMs can technically handle long sequences but degrade gracefully. By 500 tokens you've lost most early context. By 1000 tokens you're just relying on the last few hundred. If your task has 1k+ tokens of relevant context, switch to a transformer.

4. No gradient clipping

RNN training without clipnorm=1.0 is a NaN trap. Even when it doesn't NaN, the occasional exploding step degrades final performance. clipnorm=1.0 (or clipvalue=5.0) costs nothing and prevents the worst failure mode.

5. Wrong loss with from_logits

The Keras SparseCategoricalCrossentropy loss has a from_logits=True parameter. If your model's last layer has no softmax activation (recommended — softmax + log fused is more numerically stable), you must set from_logits=True. Setting it wrong silently trains the wrong objective.

6. Bidirectional in autoregressive generation

A bidirectional RNN at position $t$ has seen future tokens. Useful for classification, fatal for generation — you'd be conditioning on the answer. Always use unidirectional for autoregressive tasks.


🎯 Your Turn — Bidirectional LSTM Sentiment Classifier

Build a sentiment classifier for variable-length text sequences. The model should:

  • Take integer-encoded input of shape (batch, max_len) — padded with zeros.
  • Use an Embedding(vocab_size, embed_dim, mask_zero=True) to handle masking automatically.
  • Use one Bidirectional LSTM with 64 units (so output dim is 128 after concat).
  • Project to a single sigmoid output for binary sentiment.
  • Compile with binary cross-entropy, AdamW with clipnorm=1.0, and the accuracy metric.
python
import tensorflow as tf
from tensorflow.keras import layers, Input, Model

VOCAB_SIZE = 10_000
EMBED_DIM  = 128
MAX_LEN    = 200

def build_sentiment_model():
    inputs = Input(shape=(MAX_LEN,), dtype="int32")
    # TODO 1: Embedding with mask_zero=True
    # TODO 2: Bidirectional LSTM(64) — return_sequences=False for many-to-one
    # TODO 3: Dense(1, activation="sigmoid")
    # TODO 4: build the Model and compile it
    ...
    return model

model = build_sentiment_model()
model.summary()    # expect ~1.3M params, mostly the embedding

# Smoke test with random data
import numpy as np
X = np.random.randint(1, VOCAB_SIZE, size=(32, MAX_LEN))    # 0 is mask
X[:, 150:] = 0                                              # pad second half
y = np.random.randint(0, 2, size=(32,))
model.fit(X, y, epochs=1, batch_size=8, verbose=0)
print(model.predict(X[:3]).shape)    # (3, 1)

Run in Colab or locally with pip install tensorflow. Expected outputs in comments.

Hint 1 — Where the mask comes from layers.Embedding(VOCAB_SIZE, EMBED_DIM, mask_zero=True) emits a mask wherever the input is exactly 0. The Bidirectional LSTM that consumes its output automatically respects this mask — padded positions don't update the hidden state. You don't have to pass anything else; the mask flows through.
Hint 2 — Why return_sequences=False here Sentiment is a single label per sequence — "many-to-one". You want the LSTM's final hidden state (after the whole sequence is read), not the per-step states. return_sequences=False (the default) returns shape (batch, 128) after the bidirectional concatenation. Then a single Dense(1, sigmoid) produces the probability.
Show full solution
python
import tensorflow as tf
from tensorflow.keras import layers, Input, Model

VOCAB_SIZE = 10_000
EMBED_DIM  = 128
MAX_LEN    = 200

def build_sentiment_model():
    inputs = Input(shape=(MAX_LEN,), dtype="int32")

    # Embedding with masking — zero is the pad token, masked automatically
    x = layers.Embedding(VOCAB_SIZE, EMBED_DIM, mask_zero=True)(inputs)

    # Bidirectional LSTM — many-to-one (use only the final concatenated state)
    x = layers.Bidirectional(
        layers.LSTM(64, return_sequences=False, dropout=0.2, recurrent_dropout=0.0)
    )(x)

    # Binary sentiment head
    outputs = layers.Dense(1, activation="sigmoid")(x)

    model = Model(inputs, outputs, name="bilstm_sentiment")
    model.compile(
        optimizer=tf.keras.optimizers.AdamW(learning_rate=1e-3, clipnorm=1.0),
        loss="binary_crossentropy",
        metrics=["accuracy"],
    )
    return model


model = build_sentiment_model()
model.summary()
# Expected:
# Embedding         (None, 200, 128)    1,280,000
# Bidirectional     (None, 128)            98,816   (LSTM(64) ×2 directions)
# Dense             (None, 1)                  129
# Total params: ~1,378,945

# Smoke test
import numpy as np
X = np.random.randint(1, VOCAB_SIZE, size=(32, MAX_LEN))
X[:, 150:] = 0                                # right-pad with zeros
y = np.random.randint(0, 2, size=(32,))
model.fit(X, y, epochs=1, batch_size=8, verbose=0)
print(model.predict(X[:3]).shape)             # (3, 1)
print(model.predict(X[:3]))                   # values in (0, 1)

Design choices worth defending:

  • mask_zero=True on Embedding: makes padding invisible to the LSTM. Without it, the LSTM treats padding tokens as real input and the hidden state drifts on long-padded sequences — accuracy drops by several points.
  • Bidirectional(LSTM(64)): for sentiment, both directions matter — "not bad" vs "bad" differ on a single token whose negation is upstream. The output dim is 128 (64 forward + 64 backward concatenated).
  • dropout=0.2 on input weights, no recurrent dropout: recurrent dropout is slow on GPU (forces fallback from cuDNN to the generic implementation), and input dropout usually does enough. Set recurrent_dropout=0.0 to keep the cuDNN fast path.
  • AdamW with clipnorm=1.0: AdamW gives proper weight decay, clipnorm prevents the occasional exploding step.
  • Sigmoid + binary cross-entropy for binary classification, not softmax + categorical CE. Binary CE on a single-output sigmoid is the cheaper and more numerically stable choice.

For a real sentiment model (e.g. on IMDB), you'd add: a TextVectorization layer to tokenise raw strings, an EarlyStopping(monitor="val_accuracy", patience=3, restore_best_weights=True) callback, and possibly a second BiLSTM layer with return_sequences=True on the first. On IMDB, this architecture hits ~87-89% accuracy. A pretrained transformer fine-tuned on the same data reaches 95%+ — but it's 50× the parameters and 5× the inference latency. The LSTM is still the right choice when those constraints matter.


What You Learned

  • RNNs process sequences step by step, sharing weights across time. Hidden state is the memory.
  • Vanilla RNNs suffer vanishing/exploding gradients through BPTT. Gradient clipping (clipnorm=1.0) fixes explosion; LSTM fixes vanishing.
  • LSTM uses three gates (forget, input, output) and a cell state highway. Cell state updates by addition, so gradients flow far back through time.
  • GRU is a simplified LSTM with two gates and no separate cell state. Often equivalent performance, fewer params, faster.
  • Bidirectional RNNs read forward and backward — use when you have the whole sequence (classification, tagging). Never for streaming or autoregressive generation.
  • Stacking LSTMs requires return_sequences=True on all but the final layer. This is the single most common mistake.
  • Seq2seq uses an encoder LSTM + decoder LSTM with teacher forcing during training. The fixed context bottleneck motivated attention.
  • Attention lets the decoder weight encoder states by relevance at each step. The same mechanism transformers later built on.
  • Transformers replaced RNNs in NLP because of training parallelism, direct long-range dependencies, and pretrained ecosystems. See Transformer Architecture.
  • RNNs are still right for time series forecasting on small data, streaming inference, and edge devices.
  • Masking with Embedding(mask_zero=True) handles variable-length sequences in one line.
  • Temperature sampling controls the diversity-vs-coherence trade-off in generation. Top-k and top-p extend the idea.

This closes the AI architecture track. Next up: pick a domain — diffusion models, transformer internals, RAG systems, RLHF — and go deep. Or jump back to Transformer Architecture to see why the field pivoted.