PythonMastery
intermediate 18 min read · lesson 1 of 6 in Generative AI & LLMs

How Large Language Models Work

1 · The lesson

read

Strip away the marketing and a large language model is a single mathematical object: a function that takes a sequence of tokens and returns a probability distribution over the next token. Everything else — conversation, code generation, summarisation, tool use — is that one operation, run in a loop, on top of an enormous pile of learned statistics.

Once you internalise that, the rest of the LLM stack stops feeling like magic and starts feeling like engineering. You can reason about why models hallucinate, why prompts matter so much, why context windows are finite, and why a 70B-parameter model costs what it does to run. This lesson covers the mental model, the three training stages that produce a modern assistant, the capability landscape as of 2026, and your first programmatic API call.

Run locally with pip install anthropic and ANTHROPIC_API_KEY set. LLM APIs require keys and network access, so they do not run in the in-browser sandbox. Expected output is shown in comments.


1. An LLM Is P(next_token | previous_tokens)

Forget "AI" for a moment. The core object is a probability distribution:

python
P(next_token | token_1, token_2, ..., token_n)
+ setup added so this can run · defines P, token_2, token_n, next_token, token_1
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def P(*_a, **_kw):
    print('-> P() called')
    return _AutoMock('P()')
token_2 = _AutoMock('token_2')
token_n = _AutoMock('token_n')
next_token = _AutoMock('next_token')
token_1 = _AutoMock('token_1')

Given everything seen so far, what is the probability of every possible next token? The model outputs one number per token in its vocabulary (typically 50K-200K tokens). Pick one — usually by sampling — append it, and repeat. That is the entire generation loop.

python
# Pseudocode for the generation loop every modern LLM runs
tokens = tokenise(prompt)
while not done:
    logits = model(tokens)              # shape: [vocab_size]
    probs  = softmax(logits)
    next_tok = sample(probs)            # temperature, top-p, top-k live here
    tokens.append(next_tok)
    if next_tok == END_OF_TEXT: break
+ setup added so this can run · defines tokenise, prompt, done, model, softmax, sample, END_OF_TEXT
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def tokenise(*_a, **_kw):
    print('-> tokenise() called')
    return _AutoMock('tokenise()')
prompt = _AutoMock('prompt')
done = 1
def model(*_a, **_kw):
    print('-> model() called')
    return _AutoMock('model()')
def softmax(*_a, **_kw):
    print('-> softmax() called')
    return _AutoMock('softmax()')
def sample(*_a, **_kw):
    print('-> sample() called')
    return _AutoMock('sample()')
END_OF_TEXT = _AutoMock('END_OF_TEXT')

The model itself — billions of parameters of matrices and non-linearities — is a learned approximation of that conditional distribution. Everything interesting (reasoning, code, dialogue) emerges from how good the approximation is on text it has never seen.

A token is roughly a sub-word — "hello" is one token, "unimaginable" might be three. English averages around four characters per token. Pricing, context limits, and rate limits all count tokens, not words. See tokenisation from the NLP lesson.


2. The Training Corpus — Where the Statistics Come From

A frontier model in 2026 is trained on something like 15-30 trillion tokens. The blend (no provider publishes exact numbers, but the rough shape is public):

  • Web crawl (Common Crawl, filtered) — most of the bulk
  • Books, papers, code (GitHub, arXiv, Stack Exchange)
  • Curated high-quality text — Wikipedia, textbooks, reference works
  • Synthetic data — model-generated text used to fill gaps, increasingly common
  • Licensed data — news, proprietary corpora, paid partnerships

The model never "looks anything up". Whatever it knows is compressed into its weights during pretraining. Once training stops, the knowledge is frozen — this is the training cutoff.


3. The Three Training Stages

A raw-pretrained model can finish your sentences but cannot follow an instruction. Producing a useful assistant takes three stages:

StageWhat it doesDataCompute
PretrainingLearn next-token prediction on raw textTrillions of tokens, mostly unlabelledMillions of GPU-hours
Instruction tuning (SFT)Learn to follow prompts~10K-1M curated (prompt, response) pairsDays to weeks
RLHF / RLAIFLearn human (or AI) preferences over outputsPairs of (response_A, response_B, which_is_better)Weeks

Pretraining is where the bulk of the capability comes from. Predicting the next token on a substantial chunk of the internet forces the model to learn grammar, facts, programming languages, basic reasoning patterns, multilingual transfer — everything that lets it be useful at all.

Instruction tuning (or supervised fine-tuning, SFT) re-aims the model. After pretraining, asked "what is 2+2?", it might continue with "Solve the following equation: 3x+5=14" because that pattern occurs in textbooks. SFT teaches it that a question deserves an answer.

RLHF (Reinforcement Learning from Human Feedback) — humans (or another model, "RLAIF") rate which of two responses is better; the model is fine-tuned to maximise that preference. This is where helpfulness, honesty, and refusal behaviour are shaped. It is also where most of the model's "personality" lives.


4. Scale — The Numbers That Matter

Three axes of scale, all measured in different units:

Axis2018 (BERT)2020 (GPT-3)2026 (frontier)
Parameters340M175B~1-5T (dense or MoE)
Training tokens~3B~300B~10-30T
Context window5122K200K-2M
Training cost$thousands~$5M$100M-$1B

A parameter is a single floating-point number in the model's weights. More parameters = more capacity to memorise patterns and chain them together. Modern Mixture-of-Experts (MoE) models have trillions of total parameters but only activate a fraction per token, so inference cost scales with the active count, not the total.

The context window is the maximum number of tokens the model can attend to at once. Larger windows let you stuff in long documents, codebases, conversation history. They cost compute roughly quadratically — see the transformer lesson on why.


5. Emergent Abilities — And the Honest Footnote

The "emergent abilities" claim (Wei et al. 2022) was that certain capabilities — multi-step arithmetic, instruction following, chain-of-thought reasoning — appear suddenly once a model crosses a size threshold, rather than improving smoothly. This generated enormous excitement.

The honest 2026 view: a follow-up paper (Schaeffer et al. 2023) showed that "emergence" often disappears when you use a smoother metric. A task scored as "completely wrong / completely right" looks discontinuous; the same task scored on partial credit improves steadily with scale. The capabilities are real, but the suddenness was partly an artefact of how we measured.

Practical takeaway: bigger models are reliably better, but no specific size unlocks "general intelligence". Don't pay for Opus when Haiku does the job; don't expect a 7B local model to match a frontier API.


6. Capabilities in 2026

A frontier LLM today can plausibly:

  • Write idiomatic code in any mainstream language, given a clear spec
  • Reason through multi-step problems if prompted to think step by step
  • Use tools (calling functions you describe) and chain them together
  • Read images, charts, and short videos (multi-modal frontier models)
  • Maintain coherent conversation across 100K+ tokens of context
  • Translate between most language pairs at near-professional quality
  • Summarise, restructure, and critique long-form text

What it still cannot reliably do:

  • Know recent events beyond its training cutoff (without a tool to look them up)
  • Do exact arithmetic beyond a few digits (use a calculator tool)
  • Reason about genuinely novel problems outside its training distribution
  • Stay calibrated about its own uncertainty — it will sound confident when wrong
  • Remember anything across sessions unless you give it a memory system

7. Hallucinations — Why They Happen

Hallucination is the model generating fluent, plausible text that is factually wrong. It is not a bug to be patched; it is a direct consequence of how the model works.

The model picks the most probable next token given the prompt. If you ask "What was Hedy's PhD thesis about?", the training data contains no facts about Hedy, but it contains millions of patterns of "X's PhD thesis was about Y". The model fills in something plausible — because nothing in its training signal taught it to say "I don't know" when the answer is absent.

Mitigations:


  • RAG (next-but-one lesson) — give the model the facts in-context

  • Tool use — let it call a calculator, a search engine, a database

  • Prompt design — explicitly tell it "say 'I don't know' if uncertain"

  • Verification — check the output programmatically when accuracy matters

Hallucination cannot be eliminated, only constrained. Designing around that fact is half of building production LLM systems.


8. The Two-API World

You will use LLMs from one of two ecosystems:

Frontier (proprietary)Open-weights
PlayersAnthropic, OpenAI, Google, xAIMeta (Llama), Mistral, Qwen, DeepSeek
AccessHosted API onlyDownload weights, run anywhere
Top qualityBetter at the frontierCatching up fast, ~6-12 months behind
CostPer-tokenYour GPU, your electricity
PrivacyData leaves your infraFully local possible
Fine-tuningLimited / API-onlyFull control

For most application work, a frontier API is the right default — better quality per dollar at low-to-medium volume, no infrastructure. For high-volume, privacy-sensitive, or heavily fine-tuned workloads, open-weights wins. Plenty of teams run both: frontier for quality-critical paths, open-weights for cheap bulk.

This lesson series uses Anthropic's Claude API as the canonical example. The patterns translate directly to OpenAI's, Google's Gemini, and any OpenAI-compatible local server (vLLM, llama.cpp, Ollama).


9. Your First API Call

Install the SDK and set your key in an environment variable — never in code, never in git. See envconfig.

bash
pip install anthropic
export ANTHROPIC_API_KEY="sk-ant-..."     # or set in your shell profile

The minimal call:

python
# Run locally — needs ANTHROPIC_API_KEY in your environment.
from anthropic import Anthropic

client = Anthropic()        # reads ANTHROPIC_API_KEY from env

resp = client.messages.create(
    model="claude-opus-4-7",
    max_tokens=1024,
    messages=[{"role": "user", "content": "In one sentence, what is a token?"}],
)
print(resp.content[0].text)
# A token is a sub-word unit — typically a few characters — that a language
# model treats as the smallest piece of text it reads and emits.

resp is a structured object, not just a string. The text lives in resp.content[0].text because the response can contain multiple content blocks (text, tool calls, images). resp.usage.input_tokens and resp.usage.output_tokens tell you exactly what you were charged for — log these.


10. The Cost Model — Back of the Envelope

Per-token pricing, with output tokens roughly 4-5x more expensive than input across most providers (output is sequential and cannot be batched as cheaply). Anthropic's list prices as of September 2026, per million tokens:

TierExample modelInputOutput
CheapClaude Haiku 4.5$1$5
MidClaude Sonnet 5$2$10
TopClaude Opus 5 (and the claude-opus-4-7 used in this track)$5$25
FrontierClaude Fable 5.1$10$50

These numbers go stale fast — they have fallen every year so far. Check the provider's pricing page before you budget; the ratios are what to remember.

A back-of-envelope for a chatbot:


  • Average input: 1,000 tokens (system prompt + history + user message)

  • Average output: 300 tokens

  • 10,000 chats/day on Sonnet: 10_000 * (1000 * 2 + 300 * 10) / 1_000_000 = $50/day = ~$1,500/month

  • Same workload on Haiku: ~$750/month. On Opus: ~$3,750/month. On Fable: ~$7,500/month.

The right model choice is a 10x cost lever — smaller than it used to be (in 2024 the top tier cost 60x the cheapest), but still the biggest single number you control. Two more levers stack on top of it: prompt caching bills repeated input at a tenth of the price, and the batch API halves everything for work that can wait. We will return to cost optimisation in production AI applications.


Common Mistakes

1. Treating LLM output as ground truth
The model is fluent, not omniscient. For anything where correctness matters — facts, calculations, code that touches production — verify the output. Run the code, check the citation, parse the JSON. The model is a draft, not the source of truth.

2. Not handling rate limits and transient errors
API calls fail. The SDK raises anthropic.RateLimitError, anthropic.APITimeoutError, anthropic.APIConnectionError. Wrap calls in retry-with-backoff (see APIs for the general pattern) or use the SDK's built-in max_retries.

3. Ignoring token costs until the bill arrives
Long system prompts × high volume = surprise invoices. Log resp.usage on every call. Set a daily spend alert with your provider. Test at small volume before flipping on a bulk job.

4. Hardcoding API keys
A key committed to git is leaked forever, even if you git rm it five minutes later. Always read from environment variables. If you accidentally commit one, rotate it immediately — don't just delete the line.

5. Assuming the model "knows" things it does not
Training cutoff means anything after that date is a hallucination risk. Ask the model what year it thinks it is and you may get an answer two years stale. For current information, use search tools or RAG.


🎯 Your Turn — Build simple_chat

Write a function simple_chat(question, model="claude-opus-4-7") that:

  • reads the API key from the ANTHROPIC_API_KEY env var (raise a clear error if missing),
  • sends question as a single user message,
  • returns the assistant's text response as a plain string,
  • catches anthropic.APIError and re-raises as RuntimeError with a useful message,
  • prints the input/output token counts as a side observation (so the caller can see the cost shape).
python
# Run locally with `pip install anthropic` and ANTHROPIC_API_KEY set.
import os
from anthropic import Anthropic, APIError

def simple_chat(question: str, model: str = "claude-opus-4-7") -> str:
    # TODO 1: check ANTHROPIC_API_KEY is set; raise RuntimeError if not
    # TODO 2: create the Anthropic() client
    # TODO 3: call client.messages.create with model, max_tokens=1024, and the question
    # TODO 4: catch APIError and re-raise as RuntimeError("LLM call failed: ...") from e
    # TODO 5: print "used: <in_tok> in / <out_tok> out"
    # TODO 6: return resp.content[0].text
    ...

if __name__ == "__main__":
    print(simple_chat("Give me one sentence on why transformers replaced RNNs."))
Hint 1 — Env-var check os.environ.get("ANTHROPIC_API_KEY") returns None if missing. Raise RuntimeError("set ANTHROPIC_API_KEY before calling simple_chat"). The SDK will also raise its own error if you skip this, but a clear up-front message saves debugging time.
Hint 2 — Token usage The response object exposes resp.usage.input_tokens and resp.usage.output_tokens. Print them before returning so a developer running this in a notebook can see the cost shape at a glance. In production you would log them, not print.
Show full solution
python
# Run locally with `pip install anthropic` and ANTHROPIC_API_KEY set.
import os
from anthropic import Anthropic, APIError


def simple_chat(question: str, model: str = "claude-opus-4-7") -> str:
    """Send a one-shot question to Claude and return the text response."""
    if not os.environ.get("ANTHROPIC_API_KEY"):
        raise RuntimeError("set ANTHROPIC_API_KEY before calling simple_chat")

    client = Anthropic()
    try:
        resp = client.messages.create(
            model=model,
            max_tokens=1024,
            messages=[{"role": "user", "content": question}],
        )
    except APIError as e:
        raise RuntimeError(f"LLM call failed: {e}") from e

    print(f"used: {resp.usage.input_tokens} in / {resp.usage.output_tokens} out")
    return resp.content[0].text


if __name__ == "__main__":
    print(simple_chat("Give me one sentence on why transformers replaced RNNs."))

# Example output:
#   used: 19 in / 38 out
#   Transformers replaced RNNs because their self-attention mechanism processes
#   all tokens in parallel — enabling much longer effective context and
#   dramatically better scaling on modern GPUs.
+ setup added so this can run · defines
import os  # noqa: F401
os.environ.setdefault("ANTHROPIC_API_KEY", "example-anthropic-api-key")

What this gets right:

  • Fails fast on missing key — clearer than the SDK's own error, and surfaces config problems before a network call.
  • Wraps SDK errors — callers see one exception type (RuntimeError) instead of needing to import anthropic to catch errors.
  • Logs cost — printing token counts surfaces the cost dimension every LLM developer needs to keep an eye on.
  • from e preserves the original traceback for debugging.

What is missing for production:

  • Retries with backoff — RateLimitError and APITimeoutError are transient; the SDK can retry automatically (Anthropic(max_retries=4)), or wrap with tenacity.
  • Streaming — for anything user-facing, switch to client.messages.stream(...) so the first token appears in ~500ms instead of waiting for the full response. See production.
  • Conversation memory — this function is one-shot. A real chatbot maintains a messages list and appends to it across turns.

But the core shape — env-var key, structured request, wrapped error, logged usage — is the production starting point. Build on it as needs grow.


What You Learned

  • An LLM is a function P(next_token | previous_tokens). Generation is sampling from that distribution in a loop.
  • Tokens are sub-word units. Pricing, context limits, and rate limits all count tokens, not words.
  • Three training stages: pretraining (raw text, where capability comes from), instruction tuning (follow prompts), RLHF (match human preferences).
  • Scale matters along three axes: parameters, training tokens, context window. More is reliably better; no specific size unlocks general intelligence.
  • "Emergent abilities" are real but were partly a measurement artefact — capability grows smoothly, not magically.
  • Hallucinations are intrinsic to next-token prediction; you mitigate, you don't eliminate.
  • Two ecosystems: frontier APIs (Anthropic/OpenAI/Google) vs open-weights (Llama, Mistral, Qwen). Use both where they fit.
  • Output tokens cost ~4-5x more than input. Model tier choice is a 10x cost lever; caching and batching stack on top.
  • Always set keys in env vars, log token usage, wrap API errors, and retry transient failures.

Next: The Transformer Architecture — how the function P(next_token | previous_tokens) is actually implemented, with enough detail to reason about context windows, sampling, and efficiency tricks.