PythonMastery
beginner 25 min read · lesson 5 of 15 in Projects

Project: Word Frequency Counter

1 · The lesson

read

You'll build a program that reads a chunk of text — pasted text, a file, or a URL — and reports the most common words. The kind of tool a writer would use to spot overused words, or a data analyst to summarize social media posts.

What you'll practice: string methods (lower, split, strip), dicts, collections.Counter, regex basics, sorting.


Step 1 — The Three-Line Version

python
text = "the quick brown fox jumps over the lazy dog the fox is quick"

counts = {}
for word in text.split():
    counts[word] = counts.get(word, 0) + 1

print(counts)

dict.get(key, 0) is the trick: it returns 0 if the key isn't there, so we don't need a separate if word in counts check. Lookup, default, increment, store.

Output:

python
{'the': 3, 'quick': 2, 'brown': 1, 'fox': 2, 'jumps': 1, 'over': 1, 'lazy': 1, 'dog': 1, 'is': 1}


Step 2 — Use collections.Counter (the Pythonic Way)

The pattern "count things" is so common Python has a built-in for it.

python
from collections import Counter

text = "the quick brown fox jumps over the lazy dog the fox is quick"

counts = Counter(text.split())
print(counts)
print(counts.most_common(3))

Counter is a dict subclass with two superpowers:


  • It defaults to 0 for missing keys (no need for .get(k, 0))

  • .most_common(n) returns the top N items already sorted

Output:

python
Counter({'the': 3, 'quick': 2, 'fox': 2, 'brown': 1, ...})
[('the', 3), ('quick', 2), ('fox', 2)]

You'll see Counter in production data code constantly.


Step 3 — Real Text is Messy

Real text has uppercase, punctuation, contractions. text.split() thinks "Fox", "fox,", and "fox" are three different words.

python
from collections import Counter

text = "The fox is quick. The fox is brown! THE FOX..."

# Naive — wrong
print(Counter(text.split()).most_common(3))
# [('The', 1), ('fox', 1), ('is', 1)]   <-- "The" and "THE" not counted together

# Fix step 1: lowercase
text = text.lower()

# Fix step 2: strip punctuation. Two options:
import string

# Option A: str.translate to delete punctuation
table = str.maketrans("", "", string.punctuation)
cleaned = text.translate(table)
print(Counter(cleaned.split()).most_common(3))
# [('the', 3), ('fox', 3), ('is', 2)]

# Option B: regex (more flexible)
import re
words = re.findall(r"[a-zA-Z']+", text)
print(Counter(words).most_common(3))
# [('the', 3), ('fox', 3), ('is', 2)]

The regex [a-zA-Z']+ matches sequences of letters and apostrophes — perfect for words including contractions like "don't".


Step 4 — Remove Common "Stop Words"

The 10 most common English words ("the", "a", "is", "of", ...) are usually not interesting. Filtering them out gives you the words that actually characterize the text.

python
from collections import Counter
import re

STOP_WORDS = {
    "the", "a", "an", "and", "or", "but", "in", "on", "at", "to", "for",
    "of", "with", "by", "from", "as", "is", "was", "are", "were", "be",
    "been", "being", "have", "has", "had", "do", "does", "did", "will",
    "would", "could", "should", "may", "might", "must", "can", "this",
    "that", "these", "those", "i", "you", "he", "she", "it", "we", "they",
    "what", "which", "who", "whom", "this", "that", "am", "if", "not",
    "no", "yes", "so", "than", "then", "there", "their", "them", "its",
}

text = """
Python is a high-level, general-purpose programming language. Its design
philosophy emphasizes code readability with the use of significant indentation.
Python is dynamically typed and garbage-collected. It supports multiple
programming paradigms, including structured (particularly procedural),
object-oriented and functional programming.
"""

words = re.findall(r"[a-zA-Z']+", text.lower())
meaningful = [w for w in words if w not in STOP_WORDS and len(w) > 1]

counts = Counter(meaningful)
for word, n in counts.most_common(5):
    print(f"  {word:<15} {n}")

Output:

python
  python          2
  programming     2
  design          1
  philosophy      1
  emphasizes      1

Filtering stop words is the simplest form of text analysis — and the start of everything more advanced in NLP.


Step 5 — Read from a File

Now make it work on real files.

python
from collections import Counter
import re
import io

# In real code:
#   with open("article.txt", "r", encoding="utf-8") as f:
#       text = f.read()
# In the browser, simulate with a string:
text = """
The quick brown fox jumps over the lazy dog. The dog barks at the fox.
The fox runs away. Foxes are quick; dogs are loyal.
"""

STOP_WORDS = {"the", "a", "an", "is", "are", "at", "over", "and", "of"}

words = re.findall(r"[a-zA-Z']+", text.lower())
filtered = [w for w in words if w not in STOP_WORDS and len(w) > 2]
counts = Counter(filtered)

print(f"Total words: {len(words)}")
print(f"Unique words: {len(set(words))}")
print(f"\nTop 5 (after filtering):")
for word, n in counts.most_common(5):
    bar = "█" * n
    print(f"  {word:<10} {bar} {n}")

The little ASCII bar chart ("█" * n) is a nice touch — visual frequency at a glance.


Step 6 — Comparison Mode

Compare two texts. Which words appear in both? Which are unique to each?

python
from collections import Counter
import re

def words_of(text):
    return re.findall(r"[a-zA-Z']+", text.lower())

text_a = "the cat sat on the mat the cat was happy"
text_b = "the dog sat in the sun the dog was content"

counts_a = Counter(words_of(text_a))
counts_b = Counter(words_of(text_b))

# Counter supports set-like operations
shared = counts_a & counts_b           # min of each count
only_a = counts_a - counts_b           # items more frequent in A
only_b = counts_b - counts_a

print("Shared:", shared.most_common())
print("Only in A:", only_a.most_common())
print("Only in B:", only_b.most_common())

Counter supports +, -, & (min), | (max) — they all do what you'd expect for counts.


Step 7 — A Polished Final Version

python
from collections import Counter
import re

STOP_WORDS = {
    "the", "a", "an", "and", "or", "but", "in", "on", "at", "to", "for",
    "of", "with", "by", "from", "as", "is", "was", "are", "were", "be",
    "have", "has", "had", "i", "you", "he", "she", "it", "we", "they",
    "this", "that", "these", "those", "if", "not", "so", "than", "then",
    "there", "their", "them", "its", "am", "do", "does", "did", "will",
}

def analyze_text(text, *, min_word_length=2, top_n=10, remove_stop_words=True):
    """Return a sorted list of (word, count) tuples."""
    words = re.findall(r"[a-zA-Z']+", text.lower())
    if remove_stop_words:
        words = [w for w in words if w not in STOP_WORDS]
    words = [w for w in words if len(w) >= min_word_length]
    return Counter(words).most_common(top_n)

def print_report(text, label="Text"):
    print(f"\n=== {label} ===")
    results = analyze_text(text, top_n=8)
    max_count = max((c for _, c in results), default=1)
    for word, count in results:
        bar = "█" * int(20 * count / max_count)
        print(f"  {word:<14} {bar} {count}")

sample = """
Python is a high-level, general-purpose programming language. Its design philosophy emphasizes
code readability with significant indentation. Python is dynamically typed and garbage-collected.
"""

print_report(sample, "Python description")

Stretch Goals

1. Multiple files: aggregate counts across several files in a folder.
2. N-grams: count two-word phrases, not just single words ("quick brown", "the cat").
3. Sentence length stats: average words per sentence, longest sentence, etc.
4. Word cloud data: output JSON in the shape word-cloud libraries expect.
5. Live monitoring: watch a log file and update the counter as new lines come in (use time.sleep and seek).
6. CLI: accept a filename argument with sys.argv or argparse.


🎯 Your Turn — N-Grams (Two-Word Phrases)

Single-word frequency is useful. Pair-frequency is more interesting — it surfaces actual phrases the text uses repeatedly.

Build a function that counts the top N two-word phrases (bigrams) in a text.

python
from collections import Counter
import re

def top_bigrams(text, top_n=5):
    """Return the top N two-word phrases in the text.

    Example: 'the cat sat on the mat' has bigrams:
        ('the','cat'), ('cat','sat'), ('sat','on'), ('on','the'), ('the','mat')
    """
    # TODO 1: tokenize to lowercase words
    # TODO 2: build a list of (word_i, word_{i+1}) tuples for i = 0..len-2
    # TODO 3: feed to Counter and return .most_common(top_n)
    pass

# Test
text = "the cat sat on the mat the cat was happy the cat sat again"
print(top_bigrams(text, 3))
# Expected ~ [(('the', 'cat'), 3), (('cat', 'sat'), 2), ...]
Hint 1 — Generating pairs from a list If words = ['a', 'b', 'c', 'd'], then list(zip(words, words[1:])) gives [('a','b'), ('b','c'), ('c','d')]. The zip-with-shifted-self trick.
Hint 2 — Feeding tuples to Counter Counter works on any iterable. Pass it a list of tuples and you get tuple → count.
Show full solution
python
from collections import Counter
import re

def top_bigrams(text, top_n=5):
    words = re.findall(r"[a-zA-Z']+", text.lower())
    if len(words) < 2:
        return []
    bigrams = list(zip(words, words[1:]))
    return Counter(bigrams).most_common(top_n)

text = "the cat sat on the mat the cat was happy the cat sat again"
for (a, b), n in top_bigrams(text, 5):
    print(f"  '{a} {b}': {n}")

This is the foundation of n-gram language models — going from this to character generation is just "pick the next word by probability".


What You Learned

  • dict.get(key, default) for default values
  • collections.Counter — the Pythonic way to count things
  • String cleaning: .lower(), .translate(), string.punctuation
  • Regex re.findall(r"[a-zA-Z']+", text) for word extraction
  • List comprehensions for filtering ([w for w in words if condition])
  • Counter arithmetic (&, -) for comparison

Counter shows up everywhere in real Python. The moment you say "count occurrences of X", reach for it.

Next: Mad Libs Generator — fun with strings and templates.

Practice this

on practicepython.in

Short exercises that run in your browser and tell you what your code actually did, not just whether a test passed.