PythonMastery
advanced 30 min read · lesson 6 of 9 in AI & Deep Learning

Natural Language Processing: The Modern Stack

1 · The lesson

read

Runtime note — Hugging Face Transformers, sentence-transformers, FAISS, and spaCy don't run in the Pyodide sandbox. Execute these snippets in a local Python environment with pip install transformers sentence-transformers faiss-cpu spacy. CPU is fine for everything in this lesson except the BERT fine-tune.

In 2018, NLP was a mess of bespoke pipelines — tokenisers stitched to handwritten rules, fed into LSTMs trained from scratch on your tiny labelled corpus. In 2026, NLP is mostly transfer learning (transfer) plus a handful of well-supported libraries. The skill that matters has shifted from "implement an attention mechanism" to "pick the right pretrained model and wire it into your product".

This lesson is the practical map: the task taxonomy, the tokenisation choices, the Hugging Face ecosystem, the supporting libraries (spaCy, sentence-transformers), and the failure modes. It builds on dl-nlp, which covered the architecture side.


1. The NLP Task Taxonomy

Almost every NLP product collapses onto one of six core task shapes. Recognising which one you have is half the work.

Task familyExamplesOutput shape
ClassificationSpam vs ham, sentiment, topicOne label (or scores) per input
Token taggingNER (people, places, dates), POS, slot fillingOne label per token
GenerationTranslation, summarisation, dialog, code generationA new sequence
RetrievalSearch, semantic similarity, RAG context fetchTop-k documents by score
Question answeringExtractive (find span), generative (write answer)A span or an answer string
Structured extractionInvoice fields, function-call arguments, JSON from textA typed object

The model architecture follows from the task shape. Encoder-only models (BERT, RoBERTa) suit classification, tagging, and retrieval. Decoder-only models (GPT, Llama) suit generation and structured extraction. Encoder-decoder models (T5, BART) suit translation and summarisation. The lines have blurred since instruction-tuned LLMs took over generation, but the shape-to-architecture intuition still helps.


2. Preprocessing — A Brief History

The progression mirrors the rest of deep learning: hand-crafted rules → learned representations → bigger learned representations.

  • Rule-based (pre-2013). Tokenise on whitespace, strip stopwords, stem with Porter, hand-curate feature lists. Brittle. Language-specific.
  • Static embeddings (2013–2017). Word2Vec and GloVe — every word gets one vector, learned from co-occurrence. "Bank" gets the same vector whether it's a river bank or a financial one.
  • Contextual embeddings (2018–2020). ELMo, then BERT — the vector depends on the sentence. "Bank" near "river" looks different from "bank" near "deposit".
  • Instruction-tuned LLMs (2021–today). Forget producing a vector at all; produce text. The "representation" is a hidden state inside a model that you mostly don't introspect.

You still preprocess in 2026 — lowercasing, deduplication, PII removal — but you no longer hand-design features. The tokenizer that ships with your model decides almost everything.


3. Tokenisation — The Layer Where Bugs Live

Tokenisers split text into integer IDs. Different model families use different schemes, and using the wrong tokeniser silently produces garbage outputs.

SchemeUsed byVocabulary sizeNotes
WordClassic NLP, some LSTMs~50 k+Out-of-vocab problem — "blockchain" wasn't in 2010 vocabs
CharacterOlder char-RNNs~100No OOV, but sequences become very long
BPE (Byte Pair Encoding)GPT-2/3/4, Llama~50 kMerges frequent pairs of characters into subwords
WordPieceBERT, DistilBERT~30 kLike BPE; uses likelihood-based merges
SentencePiece (Unigram)T5, mBART, XLNet~32 kLanguage-agnostic; doesn't assume whitespace separates words

Subword tokenisers — BPE, WordPiece, SentencePiece — share one elegant property: every input can be represented losslessly, even words the model never saw during training. "Crypto-currency-mania" becomes a sequence of known subwords. That's why all modern models use them.

The practical rule: always load the tokenizer that ships with the model, and use it for both training and inference. AutoTokenizer.from_pretrained(model_name) is what makes this safe. Don't reuse a BERT tokenizer for a T5 model.


4. The Hugging Face Ecosystem

Hugging Face's transformers library is the de-facto standard for applied NLP. It exposes thousands of pretrained models behind a stable API. Two entry points cover ninety-five percent of usage:

  • pipeline() — three-line inference for a named task. The fast path.
  • AutoModel + AutoTokenizer — full control for fine-tuning or non-standard workflows.

pipeline() — The Three-Line Path

python
from transformers import pipeline

# Sentiment — defaults to a small distilbert classifier
clf = pipeline("sentiment-analysis")
print(clf("This lesson is shorter than I expected."))
# [{'label': 'POSITIVE', 'score': 0.998}]

# Named entity recognition
ner = pipeline("ner", aggregation_strategy="simple")
print(ner("Grace works on PythonMastery in Bengaluru."))
# [{'entity_group': 'PER', 'word': 'Grace', ...},
#  {'entity_group': 'ORG', 'word': 'PythonMastery', ...},
#  {'entity_group': 'LOC', 'word': 'Bengaluru', ...}]

# Summarisation
summ = pipeline("summarization", model="facebook/bart-large-cnn")
print(summ("...long article...", max_length=60, min_length=20)[0]["summary_text"])

# Translation
en_fr = pipeline("translation_en_to_fr", model="Helsinki-NLP/opus-mt-en-fr")
print(en_fr("Transfer learning saves time.")[0]["translation_text"])

# Zero-shot classification — classify without training
zsc = pipeline("zero-shot-classification")
print(zsc(
    "The stock dropped 12% on weak earnings.",
    candidate_labels=["finance", "sports", "weather", "politics"],
))
# Returns ranked scores; "finance" wins.

Zero-shot classification deserves a closer look. The model is a generic natural-language-inference classifier (typically BART-MNLI). For each candidate label it computes "does this text entail the label?", picks the highest score. You get a classifier for any label set with zero training data. The accuracy ceiling is lower than a fine-tuned model, but for prototyping or rare-label tasks it's astonishing how far you can get without writing a single training loop.

AutoModel and AutoTokenizer — Custom Workflows

python
from transformers import AutoTokenizer, AutoModel
import torch

name = "bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(name)
model = AutoModel.from_pretrained(name)

inputs = tokenizer("Hello, transformers.", return_tensors="pt", padding=True)
with torch.no_grad():
    outputs = model(**inputs)

# Last hidden state — one vector per token
print(outputs.last_hidden_state.shape)   # torch.Size([1, 7, 768])
# [CLS]-token pooled vector — a sentence-level representation
print(outputs.last_hidden_state[:, 0].shape)   # torch.Size([1, 768])

This is the layer you reach for when fine-tuning, doing custom pooling, or extracting embeddings for downstream tasks.


5. Fine-Tuning BERT for Classification

The pattern from transfer applied to text:

python
from transformers import (
    AutoTokenizer, AutoModelForSequenceClassification,
    Trainer, TrainingArguments,
)
from datasets import load_dataset

raw = load_dataset("imdb")
tokenizer = AutoTokenizer.from_pretrained("distilbert-base-uncased")

def tokenize(batch):
    return tokenizer(batch["text"], truncation=True, max_length=256)

ds = raw.map(tokenize, batched=True)
model = AutoModelForSequenceClassification.from_pretrained(
    "distilbert-base-uncased", num_labels=2,
)

args = TrainingArguments(
    output_dir="imdb-distilbert",
    learning_rate=2e-5,                  # the canonical BERT-family LR
    per_device_train_batch_size=16,
    num_train_epochs=3,
    eval_strategy="epoch",
    weight_decay=0.01,
    warmup_ratio=0.1,
)

Trainer(
    model=model, args=args,
    train_dataset=ds["train"], eval_dataset=ds["test"],
    tokenizer=tokenizer,
).train()

Three epochs of distilbert-base-uncased on IMDb hits ~92% accuracy on a single GPU in under twenty minutes. The same task from scratch with an LSTM would need days and lose by five points.

max_length=256 is a deliberate truncation. BERT-family models have a 512-token cap; truncating to 256 doubles throughput at small accuracy cost. For long documents you need a long-context model (Longformer, BigBird) or a sliding-window approach — splitting the doc into overlapping chunks, scoring each, aggregating.


6. Sequence-to-Sequence — T5 and BART

Anything that maps text to text — translation, summarisation, paraphrasing, style transfer, question generation — fits the encoder-decoder shape. T5 famously framed every task this way: the input is "translate English to German: …" and the output is the translation.

python
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

name = "t5-small"
tokenizer = AutoTokenizer.from_pretrained(name)
model = AutoModelForSeq2SeqLM.from_pretrained(name)

prompt = "summarize: " + "Transfer learning reuses pretrained weights..."
ids = tokenizer(prompt, return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=40, num_beams=4)
print(tokenizer.decode(out[0], skip_special_tokens=True))

Fine-tuning works exactly like the BERT case — use AutoModelForSeq2SeqLM and Seq2SeqTrainingArguments. Modern instruction-tuned LLMs (genai-how-llms-work) have eaten most of T5/BART's old use cases for general-purpose generation, but for narrow seq2seq tasks where you have labelled pairs and want a small, fast model, T5/BART fine-tuning is still excellent.


7. Embeddings and Retrieval — sentence-transformers + FAISS

Classification answers "what is this?". Retrieval answers "what is this like?". For search, deduplication, clustering, recommendation, and the retrieval half of RAG (genai-rag), you need embeddings: a dense vector per document such that semantically similar texts have nearby vectors.

python
from sentence_transformers import SentenceTransformer
import numpy as np
import faiss

model = SentenceTransformer("all-MiniLM-L6-v2")    # 384-dim, fast

docs = [
    "Transfer learning reuses pretrained model weights.",
    "FAISS is a library for fast similarity search.",
    "Cats are independent and aloof animals.",
    "Dogs love their owners unconditionally.",
]
embeddings = model.encode(docs, normalize_embeddings=True)

# Build a FAISS index over the embeddings
index = faiss.IndexFlatIP(embeddings.shape[1])     # inner product = cosine on normalised vecs
index.add(np.asarray(embeddings, dtype="float32"))

query = model.encode(["pretrained networks"], normalize_embeddings=True)
scores, ids = index.search(np.asarray(query, dtype="float32"), k=2)
print([docs[i] for i in ids[0]])
# ['Transfer learning reuses pretrained model weights.',
#  'FAISS is a library for fast similarity search.']

all-MiniLM-L6-v2 is the workhorse sentence embedder — 384 dimensions, fast enough to embed millions of documents on a laptop, accurate enough for most retrieval tasks. For higher accuracy at the cost of speed, all-mpnet-base-v2 (768-dim) is the standard upgrade.

FAISS is overkill for fewer than a million documents — a NumPy dot product works fine. For larger corpora, FAISS handles the algorithmic engineering (IVF, HNSW, PQ) so you don't have to. This pattern — embed → index → retrieve top-k — is the entire retrieval half of retrieval-augmented generation.


8. Modern Instruction-Tuned Models

The pipeline-API examples above use task-specific fine-tuned models. The same tasks can be solved by prompting a general-purpose LLM:

text
Classify the sentiment of this review as POSITIVE or NEGATIVE.
Review: "This lesson is shorter than I expected."
Sentiment:

For most off-the-shelf classification tasks, a competent instruction-tuned LLM (genai-how-llms-work) matches or beats a small fine-tuned encoder, with zero training. The trade-offs are predictable:

  • LLM prompting — zero training, slow, expensive per call, opaque, easy to change behaviour
  • Fine-tuned encoder — needs labelled data, fast inference, cheap at scale, deterministic, hard to change behaviour

The sweet spot in 2026: prompt an LLM for low-volume or rapidly-evolving tasks; distil to a small fine-tuned encoder once the task is settled and the volume justifies the engineering.


9. spaCy — When You Want Traditional NLP

spaCy isn't a deep-learning library; it's a fast, batteries-included NLP pipeline. POS tagging, dependency parsing, NER, lemmatisation, sentence segmentation — all in one call, all running at hundreds of thousands of tokens per second on CPU.

python
import spacy
nlp = spacy.load("en_core_web_sm")
doc = nlp("Grace shipped PythonMastery on May 14, 2026.")

for ent in doc.ents:
    print(ent.text, ent.label_)
# Grace PERSON
# PythonMastery ORG
# May 14, 2026 DATE

for token in doc:
    print(token.text, token.pos_, token.dep_, token.head.text)

Reach for spaCy when you need:

  • Speed at scale — millions of documents, CPU only.
  • Linguistic structure — dependency trees, lemmas, morphology. Transformers don't expose these as a first-class output.
  • Rule-based augmentation — Matcher and PhraseMatcher for hybrid rule-plus-ML systems.
  • Production pipelines — spaCy's Language object is serialisable, versionable, and works the same in training and serving.

spaCy and Hugging Face are complementary, not competitive. spaCy's spacy-transformers integration lets you use a HF model as the encoder inside a spaCy pipeline.


Common Mistakes

1. Training a transformer from scratch on a small dataset.

You need millions of labelled examples to beat a fine-tuned pretrained model from scratch — even more to beat one with a few thousand. If your data is under a million examples, fine-tune. If it's under ten thousand, you may not even need full fine-tuning — head-only or few-shot prompting often suffices.

2. Ignoring sequence length limits.

BERT-family models cap at 512 tokens. Naïve truncation drops the second half of every long document. If documents routinely exceed the limit, either chunk-and-aggregate, use a long-context model (Longformer, ModernBERT), or switch to an encoder-decoder that handles longer inputs.

3. Wrong tokenizer at inference.

You trained with bert-base-uncased, then loaded bert-base-cased's tokenizer at inference because of a copy-paste. The integer IDs no longer mean what the model thinks they mean and outputs are noise. Pin the model + tokenizer together, and always reload both from the same from_pretrained name.

4. Treating LLM output as ground truth.

A model that produces fluent English will produce fluent plausible English even when the underlying facts are wrong. For any user-facing system, verify the model's claims — against a database, against retrieved sources (genai-rag), or by routing high-stakes decisions to a human. The "hallucination" problem isn't a bug to be patched out; it's the default behaviour of next-token prediction.

5. Embedding without normalisation.

Cosine similarity is the inner product of unit vectors. If you forget normalize_embeddings=True and use an inner-product index, you're ranking by vector magnitude as well as direction — usually wrong, and silently. Normalise both your corpus and your queries, or use L2 distance with the un-normalised vectors and an L2 index.


🎯 Your Turn — Zero-Shot Topic Classification

Use the Hugging Face pipeline("zero-shot-classification") to classify a list of news headlines into one of five topics — without training a model. Requirements:

  • Load the zero-shot pipeline (default model is fine; it's facebook/bart-large-mnli).
  • For each headline, return the top label and its score.
  • Print results as a small table.

Headlines and candidate labels are given. The full task is wiring the pipeline call and formatting the output.

Skeleton:

python
from transformers import pipeline

headlines = [
    "Federal Reserve raises rates 25 basis points amid sticky inflation",
    "Lakers defeat Celtics in overtime thriller",
    "New strain of avian flu detected in poultry farms",
    "Senate passes bipartisan infrastructure bill",
    "Hurricane Erin strengthens to Category 4 off the Atlantic coast",
]
candidate_labels = ["finance", "sports", "health", "politics", "weather"]

# TODO 1: build the zero-shot pipeline
classifier = ...

# TODO 2: loop over headlines, classify each against candidate_labels,
#         collect (headline, top_label, top_score)

# TODO 3: print as a small table — headline truncated to 50 chars,
#         label, score formatted to 3 decimals
Hint 1 — Pipeline construction pipeline("zero-shot-classification") with no model argument uses the default facebook/bart-large-mnli — a ~400 MB download on first call.
Hint 2 — Pipeline output shape Calling classifier(text, candidate_labels=labels) returns a dict like {"sequence": ..., "labels": [...], "scores": [...]} where labels are already sorted highest-score first. So the top prediction is result["labels"][0] and result["scores"][0].
Show full solution
python
from transformers import pipeline

headlines = [
    "Federal Reserve raises rates 25 basis points amid sticky inflation",
    "Lakers defeat Celtics in overtime thriller",
    "New strain of avian flu detected in poultry farms",
    "Senate passes bipartisan infrastructure bill",
    "Hurricane Erin strengthens to Category 4 off the Atlantic coast",
]
candidate_labels = ["finance", "sports", "health", "politics", "weather"]

classifier = pipeline("zero-shot-classification")

rows = []
for h in headlines:
    result = classifier(h, candidate_labels=candidate_labels)
    rows.append((h, result["labels"][0], result["scores"][0]))

print(f"{'headline':<52} {'label':<10} {'score':>6}")
print("-" * 72)
for h, label, score in rows:
    print(f"{h[:50]:<52} {label:<10} {score:>6.3f}")

Sample output:

python
headline                                             label       score
------------------------------------------------------------------------
Federal Reserve raises rates 25 basis points amid  finance     0.954
Lakers defeat Celtics in overtime thriller         sports      0.991
New strain of avian flu detected in poultry farms  health      0.973
Senate passes bipartisan infrastructure bill       politics    0.985
Hurricane Erin strengthens to Category 4 off the   weather     0.987

Five lines of real code, no training data, results that would have been state-of-the-art in 2018. The same machinery scales to any label set you can describe in English — candidate_labels=["spam", "not spam"], ["urgent", "routine", "informational"], ["complaint", "compliment", "question"]. Zero-shot is the right first move on almost any new classification problem, even if you eventually replace it with a fine-tuned model.


What You Learned

  • The NLP task taxonomy: classification, tagging, generation, retrieval, QA, structured extraction. Architecture follows task shape.
  • Preprocessing has moved from hand-crafted rules → static embeddings → contextual embeddings → instruction-tuned LLMs. You no longer hand-design features.
  • Subword tokenisers (BPE, WordPiece, SentencePiece) eliminate the OOV problem and are why every modern model uses them. Always load the tokenizer that ships with the model.
  • Hugging Face pipeline() gives you three-line inference for the standard tasks. AutoModel + AutoTokenizer is the fine-tuning layer.
  • Fine-tuning BERT-family for classification: 2e-5 LR, 3 epochs, weight_decay=0.01, warmup. This recipe just works.
  • Embeddings + FAISS is the retrieval pattern. sentence-transformers/all-MiniLM-L6-v2 for speed, all-mpnet-base-v2 for accuracy. Always normalise for cosine.
  • Zero-shot classification gets you a usable classifier with zero training data via an NLI model.
  • spaCy is complementary to transformers — reach for it for fast pipelines, linguistic structure, and production-grade rule + ML hybrids.
  • Verify LLM outputs. Fluent ≠ correct.

Next: GANs & Generative Models — adversarial training, diffusion, and the honest 2026 verdict on which generative approach has won where.