PythonMastery
beginner 22 min read · lesson 6 of 7 in Deep Learning Fundamentals

Working with Text: NLP Basics

1 · The lesson

read

A neural network only consumes numbers. Text is symbols. Every NLP system is therefore a two-step problem: turn text into numbers (preprocessing + vectorisation), then learn from those numbers (model). Master the first half and the second is barely different from any other deep learning task.

This lesson covers the classical pipeline (tokenise → vocab → embed → pool → classify), shows it end-to-end in Keras, and points at where the field has moved with transformers.

Run in Colab or with pip install tensorflow. Outputs shown inline.


1. The Two Halves of NLP

python
"the movie was great"  →  [4, 18, 7, 92]  →  embeddings  →  pooled vector  →  prediction
        text             integer IDs       dense vectors    fixed-size       e.g. positive

Half 1 — representation: tokenise, build a vocabulary, map each token to an ID, look up a learned dense vector for that ID.

Half 2 — modelling: a normal neural network operating on those dense vectors.

Most NLP bugs live in half 1.


2. Preprocessing — The Pipeline

A typical preprocessing pipeline before vectorisation:

StepWhatWhen you skip it
Lowercase"Hello" → "hello"When case carries meaning (NER, code)
Punctuation stripping"hello!" → "hello"Modern subword tokenisers handle it
Tokenisation"hello world" → ["hello", "world"]Never skip
Stopword removaldrop "the", "and", "is"When stopwords carry signal (negation!)
Stemming / Lemmatisation"running" → "run"Embedding-based models — usually skip

For modern deep learning the trend is to do less preprocessing. An embedding layer can learn that "running" and "run" are related; you don't need to hammer them together by hand. The one step you must always do is tokenisation.


3. Tokenisation — Three Flavours

GranularityTokenVocab sizeTrade-off
Word"unbelievable" is one token10k–100kBig vocab, every typo is unseen
Charactereach letter is a token~100Tiny vocab, long sequences, slow
Subword (BPE / WordPiece)"unbelievable" → "un", "##believ", "##able"30k–50kBest of both — what every modern model uses

Classical Keras pipelines use word tokenisation. Transformer models (BERT, GPT, T5) all use subword. For your first NN text classifier, word-level via TextVectorization is fine.


4. Vocabulary and Integer IDs

A vocabulary is just a dictionary mapping token strings to integer IDs:

python
vocab = {
    "[PAD]":  0,        # padding to fixed length
    "[UNK]":  1,        # out-of-vocabulary token
    "the":    2,
    "and":    3,
    "movie":  4,
    "great":  5,
    ...
}

text = "the movie was great"
ids  = [2, 4, 1, 5]     # "was" is OOV → 1

Two special IDs you almost always need: a padding token (so sequences in a batch can be the same length) and an unknown token (for words not in the training vocabulary).


5. Embeddings — Words as Dense Vectors

A single integer ID tells the model nothing about meaning. An embedding layer maps each ID to a dense learned vector — say, 64 dimensions — that does encode meaning.

python
"king"   →  [0.12, -0.47,  0.83, ...,  0.05]
"queen"  →  [0.10, -0.45,  0.81, ..., -0.02]
"taco"   →  [0.78,  0.91, -0.12, ...,  0.66]

Two important properties of trained embeddings:

  • Similar words live near each other. "king" and "queen" end up close because they appear in similar contexts during training.
  • Direction has meaning. Famous example: "king" − "man" + "woman" ≈ "queen" (in word2vec-style embeddings).

In Keras, the embedding layer is a learnable lookup table:

python
from tensorflow.keras.layers import Embedding

# input_dim = vocab size, output_dim = embedding dimension
emb = Embedding(input_dim=10000, output_dim=64)

Input shape (batch, seq_len) of integer IDs → output shape (batch, seq_len, 64).


6. The Classical Baselines — Don't Skip Them

Before you reach for a neural network, two non-DL baselines win surprisingly often:

  • Bag-of-Words — represent a document as the count of each vocabulary word. Throw away order.
  • TF-IDF — weight those counts by how rare the word is across the corpus. Common words get downweighted.

Plug either into a LogisticRegression from scikit-learn. On small text classification datasets (under ~10k documents) this often beats a small neural network. NN advantages emerge with more data and longer documents.

Always run a TF-IDF + logistic regression baseline before investing in deep models. If a tree-based model with TF-IDF gets to 88% and your CNN gets 89% with ten times the training time, the simpler model is the right answer.


7. The Keras Text Pipeline

Modern Keras gives you TextVectorization — a preprocessing layer that lives inside the model, so the same preprocessing runs at train time and inference time. No more drift between Python preprocessing scripts and production.

python
import tensorflow as tf
from tensorflow.keras import layers
from tensorflow.keras.models import Sequential

# A small toy corpus
texts = ["the movie was great", "boring and slow", "loved every minute"]

vectorize = layers.TextVectorization(max_tokens=10000,
                                     output_mode='int',
                                     output_sequence_length=100)

# Adapt = learn vocabulary from data
vectorize.adapt(texts)

# Test it
print(vectorize(["the movie was great"]))
# → tf.Tensor([[2 5 6 7 0 0 ... 0]], shape=(1, 100), dtype=int64)

adapt() walks the training texts to build the vocabulary. After that, the layer is a deterministic text → ints function.


8. A Minimal Text Classifier

python
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import (TextVectorization, Embedding,
                                     GlobalAveragePooling1D, Dense, Dropout)

VOCAB_SIZE = 10000
SEQ_LEN    = 100
EMB_DIM    = 64

vectorize = TextVectorization(max_tokens=VOCAB_SIZE,
                              output_mode='int',
                              output_sequence_length=SEQ_LEN)
vectorize.adapt(X_train_text)         # X_train_text: a list/array of strings

model = Sequential([
    vectorize,
    Embedding(VOCAB_SIZE, EMB_DIM),
    GlobalAveragePooling1D(),
    Dropout(0.3),
    Dense(64, activation='relu'),
    Dense(1,  activation='sigmoid'),
])

model.compile(optimizer='adam',
              loss='binary_crossentropy',
              metrics=['accuracy'])

model.fit(X_train_text, y_train,
          epochs=10,
          batch_size=32,
          validation_split=0.2)
# → Epoch 10/10  loss: 0.21 - accuracy: 0.92 - val_loss: 0.30 - val_accuracy: 0.87
+ setup added so this can run · defines X_train_text, y_train
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

X_train_text = _AutoMock('X_train_text')
y_train = _AutoMock('y_train')

Read it:

1. TextVectorization turns each string into a (SEQ_LEN,) vector of integer IDs.
2. Embedding turns each ID into a 64-dim vector. Output shape: (SEQ_LEN, 64).
3. GlobalAveragePooling1D averages the 64-dim vectors across the sequence, collapsing to a single (64,) vector per document.
4. Two Dense layers — a standard binary classifier on top.

Why this works: the embeddings learn to push positive-sentiment words to similar directions; averaging them gives a "sentiment direction" for the whole document; the Dense layers project that direction onto a yes/no decision.


9. Beyond Average Pooling — RNNs

Averaging throws away word order. "not good" and "good not" produce identical pooled vectors. For tasks where order matters (translation, summarisation), you need a model that processes the sequence in order.

That's the realm of RNNs / LSTMs / GRUs — networks that maintain a hidden state and update it one token at a time. Briefly:

python
from tensorflow.keras.layers import LSTM

Sequential([
    vectorize,
    Embedding(VOCAB_SIZE, EMB_DIM),
    LSTM(64),                          # processes the sequence left-to-right
    Dense(1, activation='sigmoid'),
])
+ setup added so this can run · defines Sequential, vectorize, Embedding, VOCAB_SIZE, EMB_DIM, Dense
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def Sequential(*_a, **_kw):
    print('-> Sequential() called')
    return _AutoMock('Sequential()')
vectorize = _AutoMock('vectorize')
def Embedding(*_a, **_kw):
    print('-> Embedding() called')
    return _AutoMock('Embedding()')
VOCAB_SIZE = _AutoMock('VOCAB_SIZE')
EMB_DIM = _AutoMock('EMB_DIM')
def Dense(*_a, **_kw):
    print('-> Dense() called')
    return _AutoMock('Dense()')

LSTMs were the dominant NLP architecture from ~2014 to ~2018. They still work, but they've been mostly displaced by transformers. Full coverage in rnn.


10. The Transformer Revolution — One Paragraph

In 2017 the paper "Attention Is All You Need" introduced the transformer, a model that processes a whole sequence in parallel using self-attention (each token decides how much to pay attention to every other token). Transformers scale dramatically better than RNNs and learn long-range dependencies natively. Every modern foundation model — BERT, GPT, T5, LLaMA, Claude — is a transformer. We cover them in genai-transformer.


11. Hugging Face — Off-the-Shelf Power

You will rarely train a sentiment classifier from scratch anymore. Hugging Face's transformers library wraps every important pretrained model:

python
from transformers import pipeline

sentiment = pipeline("sentiment-analysis")

sentiment("I absolutely loved this film.")
# → [{'label': 'POSITIVE', 'score': 0.9998}]

sentiment("Two hours I will never get back.")
# → [{'label': 'NEGATIVE', 'score': 0.9994}]

Three lines, state-of-the-art accuracy, no training. For 90% of real text tasks the right starting point is a pretrained transformer fine-tuned on your data — not a fresh embedding-pool-Dense model. The pipeline above is just the easiest way in.

That said: learning the pipeline in Section 8 first builds intuition for what's happening inside those black boxes.


Common Mistakes

  • Stopword removal on tasks where stopwords matter. Sentiment ("not good", "never worked") and negation tasks blow up without stopwords. Sentiment with "not" removed flips the sign of half your training examples.
  • Inconsistent preprocessing between train and inference. Classic source of "works on my laptop, fails in production". TextVectorization inside the model prevents this — use it.
  • Vocab too large for embedding dim. A vocab of 1,000,000 with 8-dim embeddings learns nothing useful. Rule of thumb: embedding_dim ≥ 4·log2(vocab_size). For 10k vocab, 32+ dimensions.
  • Sequence length too short. Truncating reviews to 50 tokens when the average review is 300 throws away most of the signal. Plot a histogram of token counts and pick a length that covers most of the distribution.
  • Skipping the TF-IDF baseline. Don't fall in love with deep models before checking whether logistic regression already solves your problem.

🎯 Your Turn — Sentiment Classifier

Assume you have train_texts (list of strings) and train_labels (0 = negative, 1 = positive). Build a binary text classifier: TextVectorization (vocab 5000, seq length 80) → Embedding(5000, 32) → GlobalAveragePooling1D → Dense(32, relu) → Dense(1, sigmoid).

Adapt the vectorizer on the training texts, compile, and fit for 10 epochs.

Skeleton:

python
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import (TextVectorization, Embedding,
                                     GlobalAveragePooling1D, Dense)

# TODO 1: create TextVectorization (5000 tokens, sequence length 80)
vectorize = ...

# TODO 2: adapt it on train_texts

model = Sequential([
    # TODO 3: vectorize
    # TODO 4: Embedding(5000, 32)
    # TODO 5: GlobalAveragePooling1D
    # TODO 6: Dense(32, relu)
    # TODO 7: Dense(1, sigmoid)
])

# TODO 8: compile (adam, binary_crossentropy, accuracy)
# TODO 9: fit for 10 epochs with validation_split=0.2
Hint 1 — TextVectorization args TextVectorization(max_tokens=5000, output_mode='int', output_sequence_length=80). Then call vectorize.adapt(train_texts).
Hint 2 — Pooling collapses the sequence GlobalAveragePooling1D turns (batch, 80, 32) into (batch, 32) by averaging across the 80-token dimension. After it, you're back to a standard Dense classifier.
Show full solution
python
from tensorflow.keras.models import Sequential
from tensorflow.keras.layers import (TextVectorization, Embedding,
                                     GlobalAveragePooling1D, Dense)

vectorize = TextVectorization(max_tokens=5000,
                              output_mode='int',
                              output_sequence_length=80)
vectorize.adapt(train_texts)

model = Sequential([
    vectorize,
    Embedding(5000, 32),
    GlobalAveragePooling1D(),
    Dense(32, activation='relu'),
    Dense(1,  activation='sigmoid'),
])

model.compile(optimizer='adam',
              loss='binary_crossentropy',
              metrics=['accuracy'])

history = model.fit(train_texts, train_labels,
                    epochs=10,
                    batch_size=32,
                    validation_split=0.2,
                    verbose=2)
# → Epoch 10/10  loss: 0.24 - accuracy: 0.91 - val_loss: 0.34 - val_accuracy: 0.86
+ setup added so this can run · defines train_texts, train_labels
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

train_texts = _AutoMock('train_texts')
train_labels = _AutoMock('train_labels')

Because the TextVectorization layer lives inside the model, deploying it is one model.save("sentiment.keras") away — preprocessing ships with the weights. We cover that in Deploy Your Neural Network.


What You Learned

  • NLP = text → numbers → model. The first arrow is where most bugs live.
  • Use subword tokenisation in production; word-level with TextVectorization is fine for first models.
  • Embeddings are learned dense vectors per token — similar words land in similar directions.
  • The minimal classifier is TextVectorization → Embedding → GlobalAveragePooling1D → Dense.
  • For order-sensitive tasks, reach for LSTM/GRU (rnn); for state-of-the-art, transformers (genai-transformer) and the Hugging Face ecosystem.
  • Always run a TF-IDF + logistic regression baseline first.

Next: Deploy Your Neural Network — turning the model on your laptop into one your users can hit over HTTPS.