Project: Word Frequency Counter
1 · The lesson
readYou'll build a program that reads a chunk of text — pasted text, a file, or a URL — and reports the most common words. The kind of tool a writer would use to spot overused words, or a data analyst to summarize social media posts.
What you'll practice: string methods (lower, split, strip), dicts, collections.Counter, regex basics, sorting.
Step 1 — The Three-Line Version
text = "the quick brown fox jumps over the lazy dog the fox is quick" counts = {} for word in text.split(): counts[word] = counts.get(word, 0) + 1 print(counts)
dict.get(key, 0) is the trick: it returns 0 if the key isn't there, so we don't need a separate if word in counts check. Lookup, default, increment, store.
Output:
{'the': 3, 'quick': 2, 'brown': 1, 'fox': 2, 'jumps': 1, 'over': 1, 'lazy': 1, 'dog': 1, 'is': 1}Step 2 — Use collections.Counter (the Pythonic Way)
The pattern "count things" is so common Python has a built-in for it.
from collections import Counter text = "the quick brown fox jumps over the lazy dog the fox is quick" counts = Counter(text.split()) print(counts) print(counts.most_common(3))
Counter is a dict subclass with two superpowers:
- It defaults to
0for missing keys (no need for.get(k, 0)) .most_common(n)returns the top N items already sorted
Output:
Counter({'the': 3, 'quick': 2, 'fox': 2, 'brown': 1, ...}) [('the', 3), ('quick', 2), ('fox', 2)]
You'll see Counter in production data code constantly.
Step 3 — Real Text is Messy
Real text has uppercase, punctuation, contractions. text.split() thinks "Fox", "fox,", and "fox" are three different words.
from collections import Counter text = "The fox is quick. The fox is brown! THE FOX..." # Naive — wrong print(Counter(text.split()).most_common(3)) # [('The', 1), ('fox', 1), ('is', 1)] <-- "The" and "THE" not counted together # Fix step 1: lowercase text = text.lower() # Fix step 2: strip punctuation. Two options: import string # Option A: str.translate to delete punctuation table = str.maketrans("", "", string.punctuation) cleaned = text.translate(table) print(Counter(cleaned.split()).most_common(3)) # [('the', 3), ('fox', 3), ('is', 2)] # Option B: regex (more flexible) import re words = re.findall(r"[a-zA-Z']+", text) print(Counter(words).most_common(3)) # [('the', 3), ('fox', 3), ('is', 2)]
The regex [a-zA-Z']+ matches sequences of letters and apostrophes — perfect for words including contractions like "don't".
Step 4 — Remove Common "Stop Words"
The 10 most common English words ("the", "a", "is", "of", ...) are usually not interesting. Filtering them out gives you the words that actually characterize the text.
from collections import Counter import re STOP_WORDS = { "the", "a", "an", "and", "or", "but", "in", "on", "at", "to", "for", "of", "with", "by", "from", "as", "is", "was", "are", "were", "be", "been", "being", "have", "has", "had", "do", "does", "did", "will", "would", "could", "should", "may", "might", "must", "can", "this", "that", "these", "those", "i", "you", "he", "she", "it", "we", "they", "what", "which", "who", "whom", "this", "that", "am", "if", "not", "no", "yes", "so", "than", "then", "there", "their", "them", "its", } text = """ Python is a high-level, general-purpose programming language. Its design philosophy emphasizes code readability with the use of significant indentation. Python is dynamically typed and garbage-collected. It supports multiple programming paradigms, including structured (particularly procedural), object-oriented and functional programming. """ words = re.findall(r"[a-zA-Z']+", text.lower()) meaningful = [w for w in words if w not in STOP_WORDS and len(w) > 1] counts = Counter(meaningful) for word, n in counts.most_common(5): print(f" {word:<15} {n}")
Output:
python 2 programming 2 design 1 philosophy 1 emphasizes 1
Filtering stop words is the simplest form of text analysis — and the start of everything more advanced in NLP.
Step 5 — Read from a File
Now make it work on real files.
from collections import Counter import re import io # In real code: # with open("article.txt", "r", encoding="utf-8") as f: # text = f.read() # In the browser, simulate with a string: text = """ The quick brown fox jumps over the lazy dog. The dog barks at the fox. The fox runs away. Foxes are quick; dogs are loyal. """ STOP_WORDS = {"the", "a", "an", "is", "are", "at", "over", "and", "of"} words = re.findall(r"[a-zA-Z']+", text.lower()) filtered = [w for w in words if w not in STOP_WORDS and len(w) > 2] counts = Counter(filtered) print(f"Total words: {len(words)}") print(f"Unique words: {len(set(words))}") print(f"\nTop 5 (after filtering):") for word, n in counts.most_common(5): bar = "█" * n print(f" {word:<10} {bar} {n}")
The little ASCII bar chart ("█" * n) is a nice touch — visual frequency at a glance.
Step 6 — Comparison Mode
Compare two texts. Which words appear in both? Which are unique to each?
from collections import Counter import re def words_of(text): return re.findall(r"[a-zA-Z']+", text.lower()) text_a = "the cat sat on the mat the cat was happy" text_b = "the dog sat in the sun the dog was content" counts_a = Counter(words_of(text_a)) counts_b = Counter(words_of(text_b)) # Counter supports set-like operations shared = counts_a & counts_b # min of each count only_a = counts_a - counts_b # items more frequent in A only_b = counts_b - counts_a print("Shared:", shared.most_common()) print("Only in A:", only_a.most_common()) print("Only in B:", only_b.most_common())
Counter supports +, -, & (min), | (max) — they all do what you'd expect for counts.
Step 7 — A Polished Final Version
from collections import Counter import re STOP_WORDS = { "the", "a", "an", "and", "or", "but", "in", "on", "at", "to", "for", "of", "with", "by", "from", "as", "is", "was", "are", "were", "be", "have", "has", "had", "i", "you", "he", "she", "it", "we", "they", "this", "that", "these", "those", "if", "not", "so", "than", "then", "there", "their", "them", "its", "am", "do", "does", "did", "will", } def analyze_text(text, *, min_word_length=2, top_n=10, remove_stop_words=True): """Return a sorted list of (word, count) tuples.""" words = re.findall(r"[a-zA-Z']+", text.lower()) if remove_stop_words: words = [w for w in words if w not in STOP_WORDS] words = [w for w in words if len(w) >= min_word_length] return Counter(words).most_common(top_n) def print_report(text, label="Text"): print(f"\n=== {label} ===") results = analyze_text(text, top_n=8) max_count = max((c for _, c in results), default=1) for word, count in results: bar = "█" * int(20 * count / max_count) print(f" {word:<14} {bar} {count}") sample = """ Python is a high-level, general-purpose programming language. Its design philosophy emphasizes code readability with significant indentation. Python is dynamically typed and garbage-collected. """ print_report(sample, "Python description")
Stretch Goals
1. Multiple files: aggregate counts across several files in a folder.
2. N-grams: count two-word phrases, not just single words ("quick brown", "the cat").
3. Sentence length stats: average words per sentence, longest sentence, etc.
4. Word cloud data: output JSON in the shape word-cloud libraries expect.
5. Live monitoring: watch a log file and update the counter as new lines come in (use time.sleep and seek).
6. CLI: accept a filename argument with sys.argv or argparse.
🎯 Your Turn — N-Grams (Two-Word Phrases)
Single-word frequency is useful. Pair-frequency is more interesting — it surfaces actual phrases the text uses repeatedly.
Build a function that counts the top N two-word phrases (bigrams) in a text.
from collections import Counter import re def top_bigrams(text, top_n=5): """Return the top N two-word phrases in the text. Example: 'the cat sat on the mat' has bigrams: ('the','cat'), ('cat','sat'), ('sat','on'), ('on','the'), ('the','mat') """ # TODO 1: tokenize to lowercase words # TODO 2: build a list of (word_i, word_{i+1}) tuples for i = 0..len-2 # TODO 3: feed to Counter and return .most_common(top_n) pass # Test text = "the cat sat on the mat the cat was happy the cat sat again" print(top_bigrams(text, 3)) # Expected ~ [(('the', 'cat'), 3), (('cat', 'sat'), 2), ...]
Hint 1 — Generating pairs from a list
Ifwords = ['a', 'b', 'c', 'd'], then list(zip(words, words[1:])) gives [('a','b'), ('b','c'), ('c','d')]. The zip-with-shifted-self trick.
Hint 2 — Feeding tuples to Counter
Counter works on any iterable. Pass it a list of tuples and you get tuple → count.
Show full solution
from collections import Counter import re def top_bigrams(text, top_n=5): words = re.findall(r"[a-zA-Z']+", text.lower()) if len(words) < 2: return [] bigrams = list(zip(words, words[1:])) return Counter(bigrams).most_common(top_n) text = "the cat sat on the mat the cat was happy the cat sat again" for (a, b), n in top_bigrams(text, 5): print(f" '{a} {b}': {n}")
This is the foundation of n-gram language models — going from this to character generation is just "pick the next word by probability".
What You Learned
dict.get(key, default)for default valuescollections.Counter— the Pythonic way to count things- String cleaning:
.lower(),.translate(),string.punctuation - Regex
re.findall(r"[a-zA-Z']+", text)for word extraction - List comprehensions for filtering (
[w for w in words if condition]) - Counter arithmetic (
&,-) for comparison
Counter shows up everywhere in real Python. The moment you say "count occurrences of X", reach for it.
Next: Mad Libs Generator — fun with strings and templates.
Practice this
on practicepython.inShort exercises that run in your browser and tell you what your code actually did, not just whether a test passed.