How Large Language Models Work
1 · The lesson
readStrip away the marketing and a large language model is a single mathematical object: a function that takes a sequence of tokens and returns a probability distribution over the next token. Everything else — conversation, code generation, summarisation, tool use — is that one operation, run in a loop, on top of an enormous pile of learned statistics.
Once you internalise that, the rest of the LLM stack stops feeling like magic and starts feeling like engineering. You can reason about why models hallucinate, why prompts matter so much, why context windows are finite, and why a 70B-parameter model costs what it does to run. This lesson covers the mental model, the three training stages that produce a modern assistant, the capability landscape as of 2026, and your first programmatic API call.
Run locally with
pip install anthropicandANTHROPIC_API_KEYset. LLM APIs require keys and network access, so they do not run in the in-browser sandbox. Expected output is shown in comments.
1. An LLM Is P(next_token | previous_tokens)
Forget "AI" for a moment. The core object is a probability distribution:
P(next_token | token_1, token_2, ..., token_n) setup added so this can run · defines P, token_2, token_n, next_token, token_1
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def P(*_a, **_kw): print('-> P() called') return _AutoMock('P()') token_2 = _AutoMock('token_2') token_n = _AutoMock('token_n') next_token = _AutoMock('next_token') token_1 = _AutoMock('token_1')
Given everything seen so far, what is the probability of every possible next token? The model outputs one number per token in its vocabulary (typically 50K-200K tokens). Pick one — usually by sampling — append it, and repeat. That is the entire generation loop.
# Pseudocode for the generation loop every modern LLM runs tokens = tokenise(prompt) while not done: logits = model(tokens) # shape: [vocab_size] probs = softmax(logits) next_tok = sample(probs) # temperature, top-p, top-k live here tokens.append(next_tok) if next_tok == END_OF_TEXT: break
setup added so this can run · defines tokenise, prompt, done, model, softmax, sample, END_OF_TEXT
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def tokenise(*_a, **_kw): print('-> tokenise() called') return _AutoMock('tokenise()') prompt = _AutoMock('prompt') done = 1 def model(*_a, **_kw): print('-> model() called') return _AutoMock('model()') def softmax(*_a, **_kw): print('-> softmax() called') return _AutoMock('softmax()') def sample(*_a, **_kw): print('-> sample() called') return _AutoMock('sample()') END_OF_TEXT = _AutoMock('END_OF_TEXT')
The model itself — billions of parameters of matrices and non-linearities — is a learned approximation of that conditional distribution. Everything interesting (reasoning, code, dialogue) emerges from how good the approximation is on text it has never seen.
A token is roughly a sub-word — "hello" is one token, "unimaginable" might be three. English averages around four characters per token. Pricing, context limits, and rate limits all count tokens, not words. See tokenisation from the NLP lesson.
2. The Training Corpus — Where the Statistics Come From
A frontier model in 2026 is trained on something like 15-30 trillion tokens. The blend (no provider publishes exact numbers, but the rough shape is public):
- Web crawl (Common Crawl, filtered) — most of the bulk
- Books, papers, code (GitHub, arXiv, Stack Exchange)
- Curated high-quality text — Wikipedia, textbooks, reference works
- Synthetic data — model-generated text used to fill gaps, increasingly common
- Licensed data — news, proprietary corpora, paid partnerships
The model never "looks anything up". Whatever it knows is compressed into its weights during pretraining. Once training stops, the knowledge is frozen — this is the training cutoff.
3. The Three Training Stages
A raw-pretrained model can finish your sentences but cannot follow an instruction. Producing a useful assistant takes three stages:
| Stage | What it does | Data | Compute |
|---|---|---|---|
| Pretraining | Learn next-token prediction on raw text | Trillions of tokens, mostly unlabelled | Millions of GPU-hours |
| Instruction tuning (SFT) | Learn to follow prompts | ~10K-1M curated (prompt, response) pairs | Days to weeks |
| RLHF / RLAIF | Learn human (or AI) preferences over outputs | Pairs of (response_A, response_B, which_is_better) | Weeks |
Pretraining is where the bulk of the capability comes from. Predicting the next token on a substantial chunk of the internet forces the model to learn grammar, facts, programming languages, basic reasoning patterns, multilingual transfer — everything that lets it be useful at all.
Instruction tuning (or supervised fine-tuning, SFT) re-aims the model. After pretraining, asked "what is 2+2?", it might continue with "Solve the following equation: 3x+5=14" because that pattern occurs in textbooks. SFT teaches it that a question deserves an answer.
RLHF (Reinforcement Learning from Human Feedback) — humans (or another model, "RLAIF") rate which of two responses is better; the model is fine-tuned to maximise that preference. This is where helpfulness, honesty, and refusal behaviour are shaped. It is also where most of the model's "personality" lives.
4. Scale — The Numbers That Matter
Three axes of scale, all measured in different units:
| Axis | 2018 (BERT) | 2020 (GPT-3) | 2026 (frontier) |
|---|---|---|---|
| Parameters | 340M | 175B | ~1-5T (dense or MoE) |
| Training tokens | ~3B | ~300B | ~10-30T |
| Context window | 512 | 2K | 200K-2M |
| Training cost | $thousands | ~$5M | $100M-$1B |
A parameter is a single floating-point number in the model's weights. More parameters = more capacity to memorise patterns and chain them together. Modern Mixture-of-Experts (MoE) models have trillions of total parameters but only activate a fraction per token, so inference cost scales with the active count, not the total.
The context window is the maximum number of tokens the model can attend to at once. Larger windows let you stuff in long documents, codebases, conversation history. They cost compute roughly quadratically — see the transformer lesson on why.
5. Emergent Abilities — And the Honest Footnote
The "emergent abilities" claim (Wei et al. 2022) was that certain capabilities — multi-step arithmetic, instruction following, chain-of-thought reasoning — appear suddenly once a model crosses a size threshold, rather than improving smoothly. This generated enormous excitement.
The honest 2026 view: a follow-up paper (Schaeffer et al. 2023) showed that "emergence" often disappears when you use a smoother metric. A task scored as "completely wrong / completely right" looks discontinuous; the same task scored on partial credit improves steadily with scale. The capabilities are real, but the suddenness was partly an artefact of how we measured.
Practical takeaway: bigger models are reliably better, but no specific size unlocks "general intelligence". Don't pay for Opus when Haiku does the job; don't expect a 7B local model to match a frontier API.
6. Capabilities in 2026
A frontier LLM today can plausibly:
- Write idiomatic code in any mainstream language, given a clear spec
- Reason through multi-step problems if prompted to think step by step
- Use tools (calling functions you describe) and chain them together
- Read images, charts, and short videos (multi-modal frontier models)
- Maintain coherent conversation across 100K+ tokens of context
- Translate between most language pairs at near-professional quality
- Summarise, restructure, and critique long-form text
What it still cannot reliably do:
- Know recent events beyond its training cutoff (without a tool to look them up)
- Do exact arithmetic beyond a few digits (use a calculator tool)
- Reason about genuinely novel problems outside its training distribution
- Stay calibrated about its own uncertainty — it will sound confident when wrong
- Remember anything across sessions unless you give it a memory system
7. Hallucinations — Why They Happen
Hallucination is the model generating fluent, plausible text that is factually wrong. It is not a bug to be patched; it is a direct consequence of how the model works.
The model picks the most probable next token given the prompt. If you ask "What was Hedy's PhD thesis about?", the training data contains no facts about Hedy, but it contains millions of patterns of "X's PhD thesis was about Y". The model fills in something plausible — because nothing in its training signal taught it to say "I don't know" when the answer is absent.
Mitigations:
- RAG (next-but-one lesson) — give the model the facts in-context
- Tool use — let it call a calculator, a search engine, a database
- Prompt design — explicitly tell it "say 'I don't know' if uncertain"
- Verification — check the output programmatically when accuracy matters
Hallucination cannot be eliminated, only constrained. Designing around that fact is half of building production LLM systems.
8. The Two-API World
You will use LLMs from one of two ecosystems:
| Frontier (proprietary) | Open-weights | |
|---|---|---|
| Players | Anthropic, OpenAI, Google, xAI | Meta (Llama), Mistral, Qwen, DeepSeek |
| Access | Hosted API only | Download weights, run anywhere |
| Top quality | Better at the frontier | Catching up fast, ~6-12 months behind |
| Cost | Per-token | Your GPU, your electricity |
| Privacy | Data leaves your infra | Fully local possible |
| Fine-tuning | Limited / API-only | Full control |
For most application work, a frontier API is the right default — better quality per dollar at low-to-medium volume, no infrastructure. For high-volume, privacy-sensitive, or heavily fine-tuned workloads, open-weights wins. Plenty of teams run both: frontier for quality-critical paths, open-weights for cheap bulk.
This lesson series uses Anthropic's Claude API as the canonical example. The patterns translate directly to OpenAI's, Google's Gemini, and any OpenAI-compatible local server (vLLM, llama.cpp, Ollama).
9. Your First API Call
Install the SDK and set your key in an environment variable — never in code, never in git. See envconfig.
pip install anthropic export ANTHROPIC_API_KEY="sk-ant-..." # or set in your shell profile
The minimal call:
# Run locally — needs ANTHROPIC_API_KEY in your environment. from anthropic import Anthropic client = Anthropic() # reads ANTHROPIC_API_KEY from env resp = client.messages.create( model="claude-opus-4-7", max_tokens=1024, messages=[{"role": "user", "content": "In one sentence, what is a token?"}], ) print(resp.content[0].text) # A token is a sub-word unit — typically a few characters — that a language # model treats as the smallest piece of text it reads and emits.
resp is a structured object, not just a string. The text lives in resp.content[0].text because the response can contain multiple content blocks (text, tool calls, images). resp.usage.input_tokens and resp.usage.output_tokens tell you exactly what you were charged for — log these.
10. The Cost Model — Back of the Envelope
Per-token pricing, with output tokens roughly 4-5x more expensive than input across most providers (output is sequential and cannot be batched as cheaply). Anthropic's list prices as of September 2026, per million tokens:
| Tier | Example model | Input | Output |
|---|---|---|---|
| Cheap | Claude Haiku 4.5 | $1 | $5 |
| Mid | Claude Sonnet 5 | $2 | $10 |
| Top | Claude Opus 5 (and the claude-opus-4-7 used in this track) | $5 | $25 |
| Frontier | Claude Fable 5.1 | $10 | $50 |
These numbers go stale fast — they have fallen every year so far. Check the provider's pricing page before you budget; the ratios are what to remember.
A back-of-envelope for a chatbot:
- Average input: 1,000 tokens (system prompt + history + user message)
- Average output: 300 tokens
- 10,000 chats/day on Sonnet:
10_000 * (1000 * 2 + 300 * 10) / 1_000_000 = $50/day = ~$1,500/month - Same workload on Haiku:
~$750/month. On Opus:~$3,750/month. On Fable:~$7,500/month.
The right model choice is a 10x cost lever — smaller than it used to be (in 2024 the top tier cost 60x the cheapest), but still the biggest single number you control. Two more levers stack on top of it: prompt caching bills repeated input at a tenth of the price, and the batch API halves everything for work that can wait. We will return to cost optimisation in production AI applications.
Common Mistakes
1. Treating LLM output as ground truth
The model is fluent, not omniscient. For anything where correctness matters — facts, calculations, code that touches production — verify the output. Run the code, check the citation, parse the JSON. The model is a draft, not the source of truth.
2. Not handling rate limits and transient errors
API calls fail. The SDK raises anthropic.RateLimitError, anthropic.APITimeoutError, anthropic.APIConnectionError. Wrap calls in retry-with-backoff (see APIs for the general pattern) or use the SDK's built-in max_retries.
3. Ignoring token costs until the bill arrives
Long system prompts × high volume = surprise invoices. Log resp.usage on every call. Set a daily spend alert with your provider. Test at small volume before flipping on a bulk job.
4. Hardcoding API keys
A key committed to git is leaked forever, even if you git rm it five minutes later. Always read from environment variables. If you accidentally commit one, rotate it immediately — don't just delete the line.
5. Assuming the model "knows" things it does not
Training cutoff means anything after that date is a hallucination risk. Ask the model what year it thinks it is and you may get an answer two years stale. For current information, use search tools or RAG.
🎯 Your Turn — Build simple_chat
Write a function simple_chat(question, model="claude-opus-4-7") that:
- reads the API key from the
ANTHROPIC_API_KEYenv var (raise a clear error if missing), - sends
questionas a single user message, - returns the assistant's text response as a plain string,
- catches
anthropic.APIErrorand re-raises asRuntimeErrorwith a useful message, - prints the input/output token counts as a side observation (so the caller can see the cost shape).
# Run locally with `pip install anthropic` and ANTHROPIC_API_KEY set. import os from anthropic import Anthropic, APIError def simple_chat(question: str, model: str = "claude-opus-4-7") -> str: # TODO 1: check ANTHROPIC_API_KEY is set; raise RuntimeError if not # TODO 2: create the Anthropic() client # TODO 3: call client.messages.create with model, max_tokens=1024, and the question # TODO 4: catch APIError and re-raise as RuntimeError("LLM call failed: ...") from e # TODO 5: print "used: <in_tok> in / <out_tok> out" # TODO 6: return resp.content[0].text ... if __name__ == "__main__": print(simple_chat("Give me one sentence on why transformers replaced RNNs."))
Hint 1 — Env-var check
os.environ.get("ANTHROPIC_API_KEY") returns None if missing. Raise RuntimeError("set ANTHROPIC_API_KEY before calling simple_chat"). The SDK will also raise its own error if you skip this, but a clear up-front message saves debugging time.
Hint 2 — Token usage
The response object exposesresp.usage.input_tokens and resp.usage.output_tokens. Print them before returning so a developer running this in a notebook can see the cost shape at a glance. In production you would log them, not print.
Show full solution
# Run locally with `pip install anthropic` and ANTHROPIC_API_KEY set. import os from anthropic import Anthropic, APIError def simple_chat(question: str, model: str = "claude-opus-4-7") -> str: """Send a one-shot question to Claude and return the text response.""" if not os.environ.get("ANTHROPIC_API_KEY"): raise RuntimeError("set ANTHROPIC_API_KEY before calling simple_chat") client = Anthropic() try: resp = client.messages.create( model=model, max_tokens=1024, messages=[{"role": "user", "content": question}], ) except APIError as e: raise RuntimeError(f"LLM call failed: {e}") from e print(f"used: {resp.usage.input_tokens} in / {resp.usage.output_tokens} out") return resp.content[0].text if __name__ == "__main__": print(simple_chat("Give me one sentence on why transformers replaced RNNs.")) # Example output: # used: 19 in / 38 out # Transformers replaced RNNs because their self-attention mechanism processes # all tokens in parallel — enabling much longer effective context and # dramatically better scaling on modern GPUs.
setup added so this can run · defines
import os # noqa: F401 os.environ.setdefault("ANTHROPIC_API_KEY", "example-anthropic-api-key")
What this gets right:
- Fails fast on missing key — clearer than the SDK's own error, and surfaces config problems before a network call.
- Wraps SDK errors — callers see one exception type (
RuntimeError) instead of needing to importanthropicto catch errors. - Logs cost — printing token counts surfaces the cost dimension every LLM developer needs to keep an eye on.
from epreserves the original traceback for debugging.
What is missing for production:
- Retries with backoff —
RateLimitErrorandAPITimeoutErrorare transient; the SDK can retry automatically (Anthropic(max_retries=4)), or wrap withtenacity. - Streaming — for anything user-facing, switch to
client.messages.stream(...)so the first token appears in ~500ms instead of waiting for the full response. See production. - Conversation memory — this function is one-shot. A real chatbot maintains a
messageslist and appends to it across turns.
But the core shape — env-var key, structured request, wrapped error, logged usage — is the production starting point. Build on it as needs grow.
What You Learned
- An LLM is a function
P(next_token | previous_tokens). Generation is sampling from that distribution in a loop. - Tokens are sub-word units. Pricing, context limits, and rate limits all count tokens, not words.
- Three training stages: pretraining (raw text, where capability comes from), instruction tuning (follow prompts), RLHF (match human preferences).
- Scale matters along three axes: parameters, training tokens, context window. More is reliably better; no specific size unlocks general intelligence.
- "Emergent abilities" are real but were partly a measurement artefact — capability grows smoothly, not magically.
- Hallucinations are intrinsic to next-token prediction; you mitigate, you don't eliminate.
- Two ecosystems: frontier APIs (Anthropic/OpenAI/Google) vs open-weights (Llama, Mistral, Qwen). Use both where they fit.
- Output tokens cost ~4-5x more than input. Model tier choice is a 10x cost lever; caching and batching stack on top.
- Always set keys in env vars, log token usage, wrap API errors, and retry transient failures.
Next: The Transformer Architecture — how the function P(next_token | previous_tokens) is actually implemented, with enough detail to reason about context windows, sampling, and efficiency tricks.