Working with Text: NLP Basics
1 · The lesson
readA neural network only consumes numbers. Text is symbols. Every NLP system is therefore a two-step problem: turn text into numbers (preprocessing + vectorisation), then learn from those numbers (model). Master the first half and the second is barely different from any other deep learning task.
This lesson covers the classical pipeline (tokenise → vocab → embed → pool → classify), shows it end-to-end in Keras, and points at where the field has moved with transformers.
Run in Colab or with
pip install tensorflow. Outputs shown inline.
1. The Two Halves of NLP
"the movie was great" → [4, 18, 7, 92] → embeddings → pooled vector → prediction text integer IDs dense vectors fixed-size e.g. positive
Half 1 — representation: tokenise, build a vocabulary, map each token to an ID, look up a learned dense vector for that ID.
Half 2 — modelling: a normal neural network operating on those dense vectors.
Most NLP bugs live in half 1.
2. Preprocessing — The Pipeline
A typical preprocessing pipeline before vectorisation:
| Step | What | When you skip it |
|---|---|---|
| Lowercase | "Hello" → "hello" | When case carries meaning (NER, code) |
| Punctuation stripping | "hello!" → "hello" | Modern subword tokenisers handle it |
| Tokenisation | "hello world" → ["hello", "world"] | Never skip |
| Stopword removal | drop "the", "and", "is" | When stopwords carry signal (negation!) |
| Stemming / Lemmatisation | "running" → "run" | Embedding-based models — usually skip |
For modern deep learning the trend is to do less preprocessing. An embedding layer can learn that "running" and "run" are related; you don't need to hammer them together by hand. The one step you must always do is tokenisation.
3. Tokenisation — Three Flavours
| Granularity | Token | Vocab size | Trade-off |
|---|---|---|---|
| Word | "unbelievable" is one token | 10k–100k | Big vocab, every typo is unseen |
| Character | each letter is a token | ~100 | Tiny vocab, long sequences, slow |
| Subword (BPE / WordPiece) | "unbelievable" → "un", "##believ", "##able" | 30k–50k | Best of both — what every modern model uses |
Classical Keras pipelines use word tokenisation. Transformer models (BERT, GPT, T5) all use subword. For your first NN text classifier, word-level via TextVectorization is fine.
4. Vocabulary and Integer IDs
A vocabulary is just a dictionary mapping token strings to integer IDs:
vocab = {
"[PAD]": 0, # padding to fixed length
"[UNK]": 1, # out-of-vocabulary token
"the": 2,
"and": 3,
"movie": 4,
"great": 5,
...
}
text = "the movie was great"
ids = [2, 4, 1, 5] # "was" is OOV → 1Two special IDs you almost always need: a padding token (so sequences in a batch can be the same length) and an unknown token (for words not in the training vocabulary).
5. Embeddings — Words as Dense Vectors
A single integer ID tells the model nothing about meaning. An embedding layer maps each ID to a dense learned vector — say, 64 dimensions — that does encode meaning.
"king" → [0.12, -0.47, 0.83, ..., 0.05] "queen" → [0.10, -0.45, 0.81, ..., -0.02] "taco" → [0.78, 0.91, -0.12, ..., 0.66]
Two important properties of trained embeddings:
- Similar words live near each other.
"king"and"queen"end up close because they appear in similar contexts during training. - Direction has meaning. Famous example:
"king" − "man" + "woman" ≈ "queen"(in word2vec-style embeddings).
In Keras, the embedding layer is a learnable lookup table:
from tensorflow.keras.layers import Embedding # input_dim = vocab size, output_dim = embedding dimension emb = Embedding(input_dim=10000, output_dim=64)
Input shape (batch, seq_len) of integer IDs → output shape (batch, seq_len, 64).
6. The Classical Baselines — Don't Skip Them
Before you reach for a neural network, two non-DL baselines win surprisingly often:
- Bag-of-Words — represent a document as the count of each vocabulary word. Throw away order.
- TF-IDF — weight those counts by how rare the word is across the corpus. Common words get downweighted.
Plug either into a LogisticRegression from scikit-learn. On small text classification datasets (under ~10k documents) this often beats a small neural network. NN advantages emerge with more data and longer documents.
Always run a TF-IDF + logistic regression baseline before investing in deep models. If a tree-based model with TF-IDF gets to 88% and your CNN gets 89% with ten times the training time, the simpler model is the right answer.
7. The Keras Text Pipeline
Modern Keras gives you TextVectorization — a preprocessing layer that lives inside the model, so the same preprocessing runs at train time and inference time. No more drift between Python preprocessing scripts and production.
import tensorflow as tf from tensorflow.keras import layers from tensorflow.keras.models import Sequential # A small toy corpus texts = ["the movie was great", "boring and slow", "loved every minute"] vectorize = layers.TextVectorization(max_tokens=10000, output_mode='int', output_sequence_length=100) # Adapt = learn vocabulary from data vectorize.adapt(texts) # Test it print(vectorize(["the movie was great"])) # → tf.Tensor([[2 5 6 7 0 0 ... 0]], shape=(1, 100), dtype=int64)
adapt() walks the training texts to build the vocabulary. After that, the layer is a deterministic text → ints function.
8. A Minimal Text Classifier
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import (TextVectorization, Embedding, GlobalAveragePooling1D, Dense, Dropout) VOCAB_SIZE = 10000 SEQ_LEN = 100 EMB_DIM = 64 vectorize = TextVectorization(max_tokens=VOCAB_SIZE, output_mode='int', output_sequence_length=SEQ_LEN) vectorize.adapt(X_train_text) # X_train_text: a list/array of strings model = Sequential([ vectorize, Embedding(VOCAB_SIZE, EMB_DIM), GlobalAveragePooling1D(), Dropout(0.3), Dense(64, activation='relu'), Dense(1, activation='sigmoid'), ]) model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy']) model.fit(X_train_text, y_train, epochs=10, batch_size=32, validation_split=0.2) # → Epoch 10/10 loss: 0.21 - accuracy: 0.92 - val_loss: 0.30 - val_accuracy: 0.87
setup added so this can run · defines X_train_text, y_train
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) X_train_text = _AutoMock('X_train_text') y_train = _AutoMock('y_train')
Read it:
1. TextVectorization turns each string into a (SEQ_LEN,) vector of integer IDs.
2. Embedding turns each ID into a 64-dim vector. Output shape: (SEQ_LEN, 64).
3. GlobalAveragePooling1D averages the 64-dim vectors across the sequence, collapsing to a single (64,) vector per document.
4. Two Dense layers — a standard binary classifier on top.
Why this works: the embeddings learn to push positive-sentiment words to similar directions; averaging them gives a "sentiment direction" for the whole document; the Dense layers project that direction onto a yes/no decision.
9. Beyond Average Pooling — RNNs
Averaging throws away word order. "not good" and "good not" produce identical pooled vectors. For tasks where order matters (translation, summarisation), you need a model that processes the sequence in order.
That's the realm of RNNs / LSTMs / GRUs — networks that maintain a hidden state and update it one token at a time. Briefly:
from tensorflow.keras.layers import LSTM Sequential([ vectorize, Embedding(VOCAB_SIZE, EMB_DIM), LSTM(64), # processes the sequence left-to-right Dense(1, activation='sigmoid'), ])
setup added so this can run · defines Sequential, vectorize, Embedding, VOCAB_SIZE, EMB_DIM, Dense
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def Sequential(*_a, **_kw): print('-> Sequential() called') return _AutoMock('Sequential()') vectorize = _AutoMock('vectorize') def Embedding(*_a, **_kw): print('-> Embedding() called') return _AutoMock('Embedding()') VOCAB_SIZE = _AutoMock('VOCAB_SIZE') EMB_DIM = _AutoMock('EMB_DIM') def Dense(*_a, **_kw): print('-> Dense() called') return _AutoMock('Dense()')
LSTMs were the dominant NLP architecture from ~2014 to ~2018. They still work, but they've been mostly displaced by transformers. Full coverage in rnn.
10. The Transformer Revolution — One Paragraph
In 2017 the paper "Attention Is All You Need" introduced the transformer, a model that processes a whole sequence in parallel using self-attention (each token decides how much to pay attention to every other token). Transformers scale dramatically better than RNNs and learn long-range dependencies natively. Every modern foundation model — BERT, GPT, T5, LLaMA, Claude — is a transformer. We cover them in genai-transformer.
11. Hugging Face — Off-the-Shelf Power
You will rarely train a sentiment classifier from scratch anymore. Hugging Face's transformers library wraps every important pretrained model:
from transformers import pipeline sentiment = pipeline("sentiment-analysis") sentiment("I absolutely loved this film.") # → [{'label': 'POSITIVE', 'score': 0.9998}] sentiment("Two hours I will never get back.") # → [{'label': 'NEGATIVE', 'score': 0.9994}]
Three lines, state-of-the-art accuracy, no training. For 90% of real text tasks the right starting point is a pretrained transformer fine-tuned on your data — not a fresh embedding-pool-Dense model. The pipeline above is just the easiest way in.
That said: learning the pipeline in Section 8 first builds intuition for what's happening inside those black boxes.
Common Mistakes
- Stopword removal on tasks where stopwords matter. Sentiment ("not good", "never worked") and negation tasks blow up without stopwords. Sentiment with
"not"removed flips the sign of half your training examples. - Inconsistent preprocessing between train and inference. Classic source of "works on my laptop, fails in production".
TextVectorizationinside the model prevents this — use it. - Vocab too large for embedding dim. A vocab of 1,000,000 with 8-dim embeddings learns nothing useful. Rule of thumb:
embedding_dim ≥ 4·log2(vocab_size). For 10k vocab, 32+ dimensions. - Sequence length too short. Truncating reviews to 50 tokens when the average review is 300 throws away most of the signal. Plot a histogram of token counts and pick a length that covers most of the distribution.
- Skipping the TF-IDF baseline. Don't fall in love with deep models before checking whether logistic regression already solves your problem.
🎯 Your Turn — Sentiment Classifier
Assume you have train_texts (list of strings) and train_labels (0 = negative, 1 = positive). Build a binary text classifier: TextVectorization (vocab 5000, seq length 80) → Embedding(5000, 32) → GlobalAveragePooling1D → Dense(32, relu) → Dense(1, sigmoid).
Adapt the vectorizer on the training texts, compile, and fit for 10 epochs.
Skeleton:
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import (TextVectorization, Embedding, GlobalAveragePooling1D, Dense) # TODO 1: create TextVectorization (5000 tokens, sequence length 80) vectorize = ... # TODO 2: adapt it on train_texts model = Sequential([ # TODO 3: vectorize # TODO 4: Embedding(5000, 32) # TODO 5: GlobalAveragePooling1D # TODO 6: Dense(32, relu) # TODO 7: Dense(1, sigmoid) ]) # TODO 8: compile (adam, binary_crossentropy, accuracy) # TODO 9: fit for 10 epochs with validation_split=0.2
Hint 1 — TextVectorization args
TextVectorization(max_tokens=5000, output_mode='int', output_sequence_length=80). Then call vectorize.adapt(train_texts).
Hint 2 — Pooling collapses the sequence
GlobalAveragePooling1D turns (batch, 80, 32) into (batch, 32) by averaging across the 80-token dimension. After it, you're back to a standard Dense classifier.
Show full solution
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import (TextVectorization, Embedding, GlobalAveragePooling1D, Dense) vectorize = TextVectorization(max_tokens=5000, output_mode='int', output_sequence_length=80) vectorize.adapt(train_texts) model = Sequential([ vectorize, Embedding(5000, 32), GlobalAveragePooling1D(), Dense(32, activation='relu'), Dense(1, activation='sigmoid'), ]) model.compile(optimizer='adam', loss='binary_crossentropy', metrics=['accuracy']) history = model.fit(train_texts, train_labels, epochs=10, batch_size=32, validation_split=0.2, verbose=2) # → Epoch 10/10 loss: 0.24 - accuracy: 0.91 - val_loss: 0.34 - val_accuracy: 0.86
setup added so this can run · defines train_texts, train_labels
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) train_texts = _AutoMock('train_texts') train_labels = _AutoMock('train_labels')
Because the TextVectorization layer lives inside the model, deploying it is one model.save("sentiment.keras") away — preprocessing ships with the weights. We cover that in Deploy Your Neural Network.
What You Learned
- NLP = text → numbers → model. The first arrow is where most bugs live.
- Use subword tokenisation in production; word-level with
TextVectorizationis fine for first models. - Embeddings are learned dense vectors per token — similar words land in similar directions.
- The minimal classifier is
TextVectorization → Embedding → GlobalAveragePooling1D → Dense. - For order-sensitive tasks, reach for LSTM/GRU (rnn); for state-of-the-art, transformers (genai-transformer) and the Hugging Face ecosystem.
- Always run a TF-IDF + logistic regression baseline first.
Next: Deploy Your Neural Network — turning the model on your laptop into one your users can hit over HTTPS.