PythonMastery
intermediate 22 min read · lesson 2 of 6 in Generative AI & LLMs

The Transformer Architecture

1 · The lesson

read

Every frontier LLM in 2026 — Claude, GPT, Gemini, Llama, Mistral — is a decoder-only transformer. The architecture was published in 2017, in a paper called "Attention Is All You Need", and after nine years of scaling experiments, training-trick research, and efficiency engineering, no one has produced anything decisively better. The same architecture that translated English to German in 2017 now writes production code, reasons over million-token contexts, and runs on every phone.

This lesson is the engineer's view of the transformer. Not "implement it from scratch" — that is its own deep dive — but enough mechanism to reason about context limits, sampling parameters, and the efficiency tricks that make modern inference fast.

Run locally with pip install anthropic and ANTHROPIC_API_KEY set. The exercise calls the API to compare sampling temperatures; expected output is shown in comments.


1. Why Transformers Replaced RNNs

Pre-2017, sequence models were RNNs (LSTMs, GRUs) — process one token at a time, carry a hidden state forward. Two killer problems:

1. Cannot parallelise training. Token t needs token t-1's hidden state, which needs t-2's, and so on. Training is inherently sequential, which wastes the parallelism that GPUs give you.
2. Information decays over distance. By the time the hidden state has compressed 500 tokens of history into one fixed-size vector, the first 100 are mush.

The transformer fixes both at once. Self-attention lets every token directly look at every other token in the sequence — no sequential bottleneck, no information decay. Training is fully parallelisable across the sequence dimension, which is what unlocked the scaling we have seen since.

The trade-off: attention is O(n²) in sequence length. We will come back to this.


2. Three Flavours — Encoder, Decoder, Encoder-Decoder

The original 2017 paper introduced an encoder-decoder model for translation. The architecture has since fragmented into three lineages:

FamilyExamplesShapeUse case
Encoder-onlyBERT, RoBERTa, ModernBERTReads full input bidirectionally, outputs vectorsClassification, embeddings, search
Decoder-onlyGPT, Claude, Llama, MistralCausal — each token attends only to the past, predicts the nextGeneration, chat, code, almost everything in 2026
Encoder-decoderT5, BART, original TransformerEncoder reads input, decoder generates output conditioned on itTranslation, summarisation; less common now

The whole generative-AI wave is decoder-only. The other two are alive and useful — BERT-style encoders are still the right tool for embeddings (RAG lesson) — but when someone says "LLM", they almost certainly mean decoder-only.

The rest of this lesson focuses on decoder-only.


3. The Attention Mechanism — Intuition

For each token in the sequence, the model asks: which of the earlier tokens are relevant to me right now, and how much should I attend to each?

Concretely, each token produces three vectors via three learned linear projections of its embedding:

  • Query (Q) — "what am I looking for?"
  • Key (K) — "what do I offer?"
  • Value (V) — "what information will I pass on if attended to?"

To compute the new representation of a token, you take its Q, dot-product it with every other token's K to get a score, softmax those scores into attention weights, then take a weighted sum of all the V vectors. The token now has a fresh representation built from everything it found relevant in the rest of the sequence.

In a decoder, the "every other token" is restricted to earlier tokens only — this is the causal mask, which is what makes the model generative. A token at position 5 can attend to positions 1-5; it cannot peek at position 6.


4. Scaled Dot-Product Attention — The Formula

The exact operation, from the 2017 paper:

python
Attention(Q, K, V) = softmax(Q K^T / sqrt(d_k)) V

Every piece, in plain English:

  • Q K^T — pairwise dot products between every query and every key. Shape: [seq_len, seq_len]. Each entry is a raw similarity score.
  • / sqrt(d_k) — divide by the square root of the key dimension. Without this, dot products grow with d_k, pushing softmax into a near-one-hot regime where gradients vanish. The scaling keeps the variance roughly constant.
  • softmax(...) — turn scores into a probability distribution over earlier tokens. Each row sums to 1.
  • ... V — weighted sum of the value vectors. Each token gets back a mixture of everyone else's V, weighted by attention.

That is one attention head. Modern models stack many of them in parallel.


5. Multi-Head Attention

One head computes one set of attention weights — one kind of relationship. The model probably needs to track several at once: syntactic dependencies, coreference, topical similarity, positional patterns. So you run many attention heads in parallel, each with its own Q, K, V projections, and concatenate the results.

A frontier model might have 64-128 attention heads per layer. Probing studies have found that individual heads learn surprisingly interpretable patterns: one head attends to the previous token, another to the syntactic head of a noun, another to recent verbs. The model picks up these "circuits" without being told to.

python
MultiHead(Q, K, V) = concat(head_1, ..., head_h) W_O

where head_i = Attention(Q W_Q_i, K W_K_i, V W_V_i)

The W_O at the end is a learned projection back to the model dimension.


6. The Full Transformer Block

A single transformer block stacks attention + a feed-forward network with residual connections and layer normalisation:

python
x = LayerNorm(x + MultiHeadAttention(x))
x = LayerNorm(x + FeedForward(x))
+ setup added so this can run · defines LayerNorm, MultiHeadAttention, FeedForward
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def LayerNorm(*_a, **_kw):
    print('-> LayerNorm() called')
    return _AutoMock('LayerNorm()')
def MultiHeadAttention(*_a, **_kw):
    print('-> MultiHeadAttention() called')
    return _AutoMock('MultiHeadAttention()')
def FeedForward(*_a, **_kw):
    print('-> FeedForward() called')
    return _AutoMock('FeedForward()')
  • Residual connections (x + ...) let gradients flow through deep stacks without vanishing. Without them, training a 100-layer transformer is intractable.
  • LayerNorm keeps activations in a stable distribution across the network. Modern variants (RMSNorm, used in Llama and most 2024+ models) are slightly faster and work equally well.
  • FeedForward is two linear layers with a non-linearity (GELU or SwiGLU). It is where most of the model's parameters live — typically 4× wider than the model dimension.

A frontier model stacks 80-120 of these blocks. Activations flow through each one and become progressively more abstract — early layers track surface patterns (tokens, syntax), middle layers track semantics, late layers track task-specific features.


7. Positional Encoding — How the Model Knows Order

Attention is permutation-invariant by default. If you shuffle the input tokens, the attention math produces the same outputs in a different order. That is a problem — "the dog bit the man" and "the man bit the dog" should not look identical.

The fix: inject position information into each token's embedding. Three approaches, in roughly chronological order:

MethodHowUsed by
Sinusoidal (2017)Add fixed sin/cos vectors of varying frequencyOriginal Transformer
Learned absoluteOne trainable embedding per positionBERT, GPT-2/3
RoPE (Rotary)Rotate Q and K in 2D pairs by position-dependent anglesLlama, Mistral, most modern LLMs
ALiBiAdd a position-decay bias to attention scoresSome open-weights models

RoPE is the de facto standard in 2026 because it has the cleanest extrapolation properties — a model trained on 8K context can be coaxed to handle 32K+ at inference time by interpolating the rotation frequencies. This is most of how context windows got so big so fast.


8. Why Long Context Is Hard — O(n²)

Attention is quadratic in sequence length. Doubling context quadruples the compute and memory for the attention operation. For a 1M-token context, the raw attention matrix would be 1 trillion entries.

Three families of fixes make modern long context possible:

FlashAttention (Dao 2022): the attention math is the same, but you reorganise the computation to never materialise the full n×n matrix in memory. Process it tile by tile, on chip. ~2-4x speed-up, much lower memory.

Sliding-window attention: each token only attends to the last W tokens (e.g. 4K). O(n·W) instead of O(n²). Pairs well with periodic "global" tokens for long-range information. Mistral uses this.

Mixture-of-Experts (MoE): orthogonal to context length, but worth mentioning. The feed-forward layer is replaced with N "expert" sub-networks; a router picks k of them per token. Total parameters scale, but per-token compute does not. DeepSeek, Mixtral, and most 2025+ frontier models are MoE.

Combined, these tricks are why a million-token context window is technically achievable on a single inference machine. They do not make the quality of attention over 1M tokens trivial — there is real ongoing research on whether models can actually use that much context effectively.


9. The KV Cache — Why Inference Is Cheap

During generation, you produce one new token at a time. Naively, every new token would re-compute attention from scratch over the whole prefix — wasteful, because for tokens 1 to t-1, the K and V vectors are exactly the same as last step.

The KV cache stores K and V for every past token. To generate token t+1, you only compute the new K_t, V_t, append them to the cache, and run attention against the whole cache. Generation becomes O(n) per token instead of O(n²).

Practical consequences:

  • The KV cache lives in GPU memory and is the dominant memory cost during inference at long contexts.
  • Prompt caching (Anthropic, OpenAI, Google all offer it) is essentially "save the KV cache to disk between requests so the same system prompt does not get re-encoded". A massive cost lever in production — see production.
  • Running a model locally without a KV cache (e.g. a buggy custom inference loop) is vastly slower. Always check.

10. The Generation Loop

Putting it together, every modern LLM runs this loop at inference time:

python
# Decoder-only inference, pseudocode
def generate(prompt: str, max_tokens: int = 256) -> str:
    tokens = tokenise(prompt)
    kv_cache = []                                  # one (K, V) entry per layer

    # Prefill: process the whole prompt in parallel, fill the KV cache
    logits, kv_cache = model.forward(tokens, kv_cache)

    output = []
    for _ in range(max_tokens):
        # Sample next token from the last position's logits
        next_tok = sample(logits[-1], temperature=0.7, top_p=0.9)
        if next_tok == END_OF_TEXT:
            break
        output.append(next_tok)

        # Decode: process just the new token, append to KV cache
        logits, kv_cache = model.forward([next_tok], kv_cache)

    return detokenise(output)
+ setup added so this can run · defines tokenise, detokenise, model, sample, END_OF_TEXT
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def tokenise(*_a, **_kw):
    print('-> tokenise() called')
    return _AutoMock('tokenise()')
def detokenise(*_a, **_kw):
    print('-> detokenise() called')
    return _AutoMock('detokenise()')
model = _AutoMock('model')
def sample(*_a, **_kw):
    print('-> sample() called')
    return _AutoMock('sample()')
END_OF_TEXT = _AutoMock('END_OF_TEXT')

Two phases, two different cost profiles:

  • Prefill (the whole prompt at once) — compute-bound, parallelisable, fast per token but pays for every input token once.
  • Decode (one token at a time) — memory-bandwidth-bound, sequential, why output tokens cost 4-5x more than input.

Streaming responses (stream=True in the SDK) start emitting tokens as soon as the prefill is done. That is why a 30-token reply feels instant and a 3,000-token essay still takes seconds — most of the wait is decode.


11. Sampling — Temperature, top-p, top-k

sample(logits, ...) is where you control creativity. The raw logits become a distribution; you choose how to draw from it.

Temperature divides the logits before softmax:

python
probs = softmax(logits / temperature)
+ setup added so this can run · defines softmax, logits, temperature
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def softmax(*_a, **_kw):
    print('-> softmax() called')
    return _AutoMock('softmax()')
logits = _AutoMock('logits')
temperature = _AutoMock('temperature')
  • T = 0 → take the argmax (deterministic, "greedy"). Best for facts, code, classification.
  • T = 0.7 → moderate randomness. Sensible default for assistants.
  • T = 1.0 → use the model's raw distribution.
  • T > 1 → flatten — more diverse, less coherent.
  • T = 2+ → noise. Rarely useful.

Top-p (nucleus) sampling: keep only the smallest set of tokens whose cumulative probability is ≥ p, renormalise, sample from that. p = 0.9 is standard. Avoids the long tail of very-unlikely-but-occasionally-catastrophic tokens.

Top-k sampling: keep only the top k tokens by probability. Cruder than top-p; rarely used alone in modern APIs.

In practice you set temperature and top_p, and leave top_k alone. For deterministic, reproducible output (testing, structured generation), use temperature=0. For creative writing, push temperature up.

Stop tokens (or stop sequences) — additional strings that, if generated, halt the loop. Useful when the model has a habit of running on after the real answer.


12. Multi-Modal — A One-Paragraph Aside

Modern frontier models are not just text. Vision transformers (ViT) chop an image into a grid of patches, treat each patch as a token, and run the same attention machinery. CLIP trained a vision encoder and a text encoder jointly so that an image's embedding and a caption's embedding land near each other in vector space — the foundation of every text-to-image and image search system. Frontier multi-modal models (Claude, GPT-4o, Gemini) fuse vision-encoder outputs with the text token stream so the LLM can "read" images natively. The architecture is the same transformer; the tokens just sometimes come from pixels.


Common Mistakes

1. Thinking attention "understands"
Attention weights tell you what the model is looking at, not why. A head that consistently attends to the previous token has not learned grammar; it has learned a useful statistic. Interpretability research is full of stories of plausible attention patterns that turn out to be nothing like the human-readable explanation. Treat attention visualisations as diagnostic, not explanatory.

2. Assuming a bigger context window helps if your prompt is poorly structured
A 1M-token context does not rescue a poorly organised prompt. Models exhibit "lost in the middle" — information in the middle of a long context is recalled less reliably than information at the beginning or end. Put critical instructions at the start and end; structure long context with clear sections (XML tags help — see prompt engineering); don't dump 500K tokens of trivia and expect the model to find the needle.

3. Not using the KV cache when running models locally
If you write your own inference loop with transformers and forget past_key_values, you re-encode the entire prefix on every generated token. A 200-token response takes 40× longer than it should. Always pass and update the KV cache, or use model.generate(...) which handles it for you.

4. Misreading temperature
temperature=0 is not "more accurate" — it is more deterministic. For some tasks (creative ideation, brainstorming) it actively hurts because the model gets stuck in repetitive grooves. temperature=1 is not "broken" — it is the model's intrinsic distribution. Pick deliberately based on the task.

5. Ignoring tokenisation when counting "characters"
"My prompt is 5,000 words, that should fit in 8K context!" — except non-English text, code, and rare tokens often hit 1-2 tokens per character. Always measure with the actual tokeniser (anthropic SDK has client.count_tokens(...)).


🎯 Your Turn — Temperature Comparison

Write generate_with_temperature(prompt, temps=(0.0, 0.7, 1.0)) that:

  • calls the Anthropic API once per temperature in temps,
  • uses the same prompt each time,
  • returns a dict mapping each temperature to the response text,
  • also prints a short side-by-side comparison.

A creative prompt makes the differences obvious. Try "Write a one-sentence opening for a noir detective story set on Mars.".

python
# Run locally with `pip install anthropic` and ANTHROPIC_API_KEY set.
from anthropic import Anthropic

def generate_with_temperature(prompt: str, temps=(0.0, 0.7, 1.0), model="claude-opus-4-7"):
    # TODO 1: create the Anthropic() client
    # TODO 2: for each temperature in temps, call client.messages.create with that temperature
    # TODO 3: store the text in a dict keyed by temperature
    # TODO 4: print "T=<t>: <first 80 chars>..." for each so we can see them side-by-side
    # TODO 5: return the dict
    ...

if __name__ == "__main__":
    out = generate_with_temperature(
        "Write a one-sentence opening for a noir detective story set on Mars."
    )
Hint 1 — Passing temperature client.messages.create accepts a temperature keyword argument (float between 0 and 1). Pass it on each call. Keep max_tokens moderate (e.g. 200) — you only need a sentence.
Hint 2 — Reading the response The text lives in resp.content[0].text. To compare side-by-side, slice it: resp.content[0].text[:80] for a preview line. Make sure your loop variable for temperature does not clash with built-in names.
Show full solution
python
# Run locally with `pip install anthropic` and ANTHROPIC_API_KEY set.
from anthropic import Anthropic


def generate_with_temperature(
    prompt: str,
    temps=(0.0, 0.7, 1.0),
    model: str = "claude-opus-4-7",
) -> dict[float, str]:
    """Call the API once per temperature and return a {temperature: text} mapping."""
    client = Anthropic()
    results: dict[float, str] = {}

    for t in temps:
        resp = client.messages.create(
            model=model,
            max_tokens=200,
            temperature=t,
            messages=[{"role": "user", "content": prompt}],
        )
        results[t] = resp.content[0].text.strip()

    print(f"\nPROMPT: {prompt}\n" + "-" * 60)
    for t, text in results.items():
        preview = text.replace("\n", " ")[:120]
        print(f"T={t}: {preview}")
    return results


if __name__ == "__main__":
    generate_with_temperature(
        "Write a one-sentence opening for a noir detective story set on Mars."
    )

# Example output (will differ each run for T > 0):
#
# PROMPT: Write a one-sentence opening for a noir detective story set on Mars.
# ------------------------------------------------------------
# T=0.0: The dust storms on Mars had a way of swallowing secrets, but they couldn't
#        hide the body slumped against the hab door of Olympus...
# T=0.7: Rain never fell on Mars, but neon did — pooling in the cracked streets of
#        Tharsis like the last neon of a dying world...
# T=1.0: She walked into my office wearing a vacuum suit two sizes too small and
#        a story three times too big, dragging red sand across the polysteel floor...

What this shows:

  • T=0 is reproducible — run it twice, you get the same sentence character-for-character. Useful for tests, classifications, deterministic structured output.
  • T=0.7 is the sweet spot for most assistant tasks — fluent, varied across runs, still coherent.
  • T=1.0 produces more lexical diversity (rare words, unusual phrasings) at the cost of occasional incoherence.

Worth doing on your own: try temperature=0 on a creative task and notice how the model gets stuck in a single "best" framing; try temperature=1.5 on a factual question and watch quality collapse. The right temperature is task-dependent.


What You Learned

  • Decoder-only transformers (Claude, GPT, Llama) are the architecture of every modern LLM.
  • Self-attention lets every token directly attend to every earlier token: softmax(Q K^T / sqrt(d_k)) V.
  • Multi-head attention runs many attention heads in parallel; each learns different relationships.
  • A transformer block: attention → add & norm → feed-forward → add & norm. Stack 80-120 of them.
  • Positional encoding (sinusoidal, learned, RoPE) tells the model token order. RoPE is the 2026 default.
  • Attention is O(n²) in sequence length — long contexts are hard. FlashAttention, sliding-window, MoE, and KV caching make modern inference fast.
  • The KV cache turns generation from O(n²) per token into O(n) per token. It is the dominant memory cost at long contexts and the basis of prompt caching.
  • Prefill vs decode — prefill is parallel and fast per token; decode is sequential and slow. This is why output tokens cost ~4-5x more.
  • Temperature controls determinism vs creativity; top-p trims the long tail. Pick deliberately.
  • Multi-modal models reuse the same architecture — images become patch tokens, fused with text tokens.

Next: Prompt Engineering — the techniques that turn a raw API call into a useful, structured, reliable LLM-powered feature.