Regular Expressions
1 · The lesson
readA regex is a tiny pattern language for describing shapes of text. "An email address." "A date in YYYY-MM-DD form." "Three digits followed by a dash." You write the shape; the engine finds (or rejects) matches.
Regex has a reputation for being write-only — and it earns it when people reach for it where they shouldn't. This lesson teaches the parts you'll actually use, when to use them, and — equally important — when to put the regex down and pick up a real parser.
1. The re Module — Your First Pattern
import re text = "Order 4271 shipped on 2026-03-14" m = re.search(r"\d+", text) # find the first run of digits print(m) # <re.Match object; span=(6, 10), match='4271'> print(m.group()) # '4271' print(m.span()) # (6, 10) — start and end indices
Three things to register from this example:
1. The r"..." raw string prefix. Always. Regex uses \ for everything (\d, \b, \w); without r, Python's string parser eats them first and you get bugs or warnings. We'll come back to this in Section 10.
2. re.search finds the first match anywhere in the string. It returns a Match object or None.
3. Match.group() gives you the matched text. Match.span() gives you the position.
2. The Five Functions You'll Use 95% of the Time
import re text = "alice@example.com, bob@example.com, carol@other.org" # 1. search — first match anywhere, or None m = re.search(r"\w+@\w+\.\w+", text) print(m.group()) # alice@example.com # 2. match — match only at the START of the string print(re.match(r"\w+@", text)) # None — text starts with "alice@example.com," but match anchors at index 0 print(re.match(r"alice", text)) # Match object # 3. findall — every match, as a list of strings print(re.findall(r"\w+@\w+\.\w+", text)) # ['alice@example.com', 'bob@example.com', 'carol@other.org'] # 4. finditer — every match, as Match objects (preferred when you need positions) for m in re.finditer(r"\w+@\w+\.\w+", text): print(m.group(), "at", m.span()) # 5. sub — find and replace print(re.sub(r"\w+@\w+\.\w+", "[EMAIL]", text)) # [EMAIL], [EMAIL], [EMAIL]
search vs match is the rookie pitfall: match is anchored at the start, search is anywhere. When in doubt, search. match is mainly useful when you're parsing a known format from position 0 (like a line in a structured log).
findall is convenient but gives you only strings — no position info, and if you use groups it returns tuples of groups (no full match). finditer is what you reach for once you outgrow findall.
3. Character Classes
The building blocks for "what kind of character can be here?"
| Pattern | Means |
|---|---|
\d | A digit. Equivalent to [0-9] for ASCII. |
\D | A non-digit. |
\w | A "word" character — letter, digit, or underscore. |
\W | A non-word character. |
\s | Whitespace — space, tab, newline. |
\S | Non-whitespace. |
. | Any character except newline (unless re.DOTALL is set). |
[abc] | Any one of a, b, or c. |
[^abc] | Any character that is NOT a, b, or c. |
[a-z] | Any lowercase letter. Ranges work in sets. |
[A-Za-z0-9_] | Hand-rolled \w. Useful when you want to be explicit. |
re.findall(r"\d", "abc123") # ['1', '2', '3'] re.findall(r"\d+", "abc123 def456") # ['123', '456'] re.findall(r"[aeiou]", "regex lesson") # ['e', 'e', 'e', 'o'] re.findall(r"[^aeiou\s]", "regex") # ['r', 'g', 'x'] — consonants
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
Inside [], most metacharacters lose their special meaning — [.] is a literal dot, [+] is a literal plus. The exceptions are ^ (negation, only at start), - (range), ] (closes the class), and \ (escape).
4. Quantifiers — How Many?
| Pattern | Means |
|---|---|
? | 0 or 1 (optional) |
* | 0 or more |
+ | 1 or more |
{n} | Exactly n |
{n,m} | Between n and m |
{n,} | At least n |
All of these are greedy by default — they grab as much as possible before backtracking. Add ? after them to make them lazy (also called non-greedy or reluctant).
text = "<b>bold</b> and <i>italic</i>" re.findall(r"<.*>", text) # ['<b>bold</b> and <i>italic</i>'] — greedy re.findall(r"<.*?>", text) # ['<b>', '</b>', '<i>', '</i>'] — lazy
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
The greedy version starts at the first <, then .* eats everything including the rest of the angle brackets, then backtracks until it finds the last >. The lazy version stops at the first > after each <. Both are valid; you almost always want lazy for tag-like delimiters.
(That said: if you're actually parsing HTML, see Section 11. Don't.)
5. Anchors — Where in the String?
| Pattern | Means |
|---|---|
^ | Start of string (or start of line with re.MULTILINE) |
$ | End of string (or end of line with re.MULTILINE) |
\b | Word boundary — the transition between \w and \W |
\B | Non-word-boundary |
The most underused of the four is \b. It matches the position between a word character and a non-word character (or string edge), with zero width. Brilliant for "match the word cat but not catalogue":
text = "cat caterpillar scattered cats" re.findall(r"cat", text) # ['cat', 'cat', 'cat', 'cat'] — too many re.findall(r"\bcat\b", text) # ['cat'] — just the word re.findall(r"\bcat", text) # ['cat', 'cat', 'cat'] — word-start cats
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
^ and $ by default match the start/end of the whole string. With re.MULTILINE, they match the start/end of each line:
log = "ERROR: x\nINFO: y\nERROR: z" re.findall(r"^ERROR", log) # ['ERROR'] — only line 1's ERROR re.findall(r"^ERROR", log, re.MULTILINE) # ['ERROR', 'ERROR']
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
6. Groups — Capturing Pieces
Parentheses (...) create a capture group — a sub-match you can pull out by index.
m = re.search(r"(\d{4})-(\d{2})-(\d{2})", "Released 2026-03-14, today.") print(m.group(0)) # 2026-03-14 — the full match print(m.group(1)) # 2026 — first group print(m.group(2)) # 03 — second group print(m.group(3)) # 14 — third group print(m.groups()) # ('2026', '03', '14')
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
Non-capturing groups (?:...) group without capturing — useful when you need grouping for alternation or quantifiers but don't want the result in .groups():
re.findall(r"(?:Mr|Mrs|Ms)\. (\w+)", "Mr. Stark, Ms. Potts") # ['Stark', 'Potts'] — only the name is captured; the title isn't
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
Named groups (?P<name>...) are the readable upgrade. You access them by name instead of fragile indices:
m = re.search(r"(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})", "2026-03-14") print(m.group("year")) # 2026 print(m.groupdict()) # {'year': '2026', 'month': '03', 'day': '14'}
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
For anything more than two groups, name them. Index 1 vs index 2 confusion has burned everyone.
7. Backreferences
Once you've captured a group, you can refer to it later — both inside the pattern (to require a repeated match) and in a replacement string (to reuse the captured text).
# In the pattern: find words that are immediately repeated re.findall(r"\b(\w+) \1\b", "the the cat sat on on the mat") # ['the', 'on'] # In a replacement: swap first and last name re.sub(r"(\w+) (\w+)", r"\2 \1", "Linus Teja") # 'Teja Linus' # Named variant re.sub(r"(?P<first>\w+) (?P<last>\w+)", r"\g<last> \g<first>", "Linus Teja") # 'Teja Linus'
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
\1 is the first captured group; \g<name> is the named-group form in replacement strings. Backreferences are how regex stops being purely "shape-matching" and starts being able to express relationships between parts of the input.
8. Flags
Flags modify how the engine interprets the pattern. Pass them as the third argument, or combine with |.
| Flag | Effect |
|---|---|
re.IGNORECASE (or re.I) | Case-insensitive matching. |
re.MULTILINE (or re.M) | ^ and $ match line boundaries, not just string. |
re.DOTALL (or re.S) | . matches newlines too. |
re.VERBOSE (or re.X) | Allow whitespace + # comments inside the pattern. |
re.findall(r"hello", "Hello world", re.IGNORECASE) # ['Hello'] re.findall(r".+", "line1\nline2", re.DOTALL) # ['line1\nline2']
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
re.VERBOSE is the one most people miss, and it turns intimidating patterns into readable ones:
phone_pattern = re.compile(r""" \+? # optional leading + (?P<country>\d{1,3}) # 1-3 digit country code [\s\-]? # optional separator (?P<area>\d{3}) # area code [\s\-]? # optional separator (?P<line>\d{4}) # line number """, re.VERBOSE) m = phone_pattern.search("Call me at +1 555-7890 maybe") print(m.groupdict()) # {'country': '1', 'area': '555', 'line': '7890'}
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
Whitespace and #-comments inside the pattern are ignored. Want to match a literal space? \ or [ ]. Any non-trivial pattern that lives in your codebase for more than a week should be in re.VERBOSE form — your future self will read it instead of squinting at it.
9. re.compile — Precompile for Hot Loops
re.search, re.findall, etc. cache the compiled pattern internally — so calling them in a loop is fine, up to a point. For very hot loops, or to give a pattern a meaningful name, compile it once and reuse:
EMAIL_RE = re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.-]+\b") for line in big_log_file: for match in EMAIL_RE.finditer(line): process(match.group())
setup added so this can run · defines big_log_file, re, process
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) big_log_file = ["alpha", "beta", "gamma"] re = _AutoMock('re') def process(*_a, **_kw): print('-> process() called') return _AutoMock('process()')
Compiled patterns expose the same operations as module-level methods — EMAIL_RE.search, .findall, .finditer, .sub. They also make patterns first-class: you can pass them around, store them in dicts, and most importantly name them so the reader of your code knows what you're matching, not just how.
10. Always Use Raw Strings — r"..."
Without r, Python's string literal parser interprets backslashes first:
# WRONG — Python sees "\d" inside a normal string. "\d" isn't a known escape, # so it stays as backslash-d, BUT this raises DeprecationWarning and will # eventually be a SyntaxError. And for things like "\b" — that IS a known # escape (backspace, \x08) — Python eats it before regex ever sees it. pattern = "\bword\b" # backspace-word-backspace, not "\b word \b" # RIGHT — raw string passes the backslashes through untouched pattern = r"\bword\b"
The fix is one character. Add r in front of every regex pattern, every time. Make it muscle memory.
The same applies to replacement strings in re.sub if they contain backslashes:
re.sub(r"(\d+)", r"<\1>", "got 42 items") # 'got <42> items'
setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) re = _AutoMock('re')
11. When NOT to Use Regex
This section saves more pain than the previous ten combined. Regex matches patterns. It does not parse structure.
| You want to… | Use… | Not regex because… |
|---|---|---|
| Parse HTML / XML | BeautifulSoup, lxml (see scraping) | HTML is nested, regex isn't. You'll get the simple case and the hard case will eat your weekend. |
| Parse JSON | json module | JSON is structured; json.loads validates and decodes in one call. |
| Parse CSV | csv module | Embedded commas, quoted fields, multi-line cells — csv handles it; regex will miss edge cases. |
| Validate emails "properly" | Send a confirmation email | The RFC 5322 grammar is ~6 pages. Any regex you find online is wrong in some case. The only true validator is "does mail get delivered?" |
| Parse Python / SQL / any code | An actual parser (ast, sqlglot) | Code is hierarchical; regex is flat. |
| Extract a number from a sentence | Regex ✓ | This is genuinely a flat pattern. Go for it. |
The rule of thumb: if the input has nested or hierarchical structure, regex is the wrong tool. If it has flat, locally-recognisable patterns, regex is excellent.
12. Real-World Patterns You'll Actually Reuse
import re # Split on any run of whitespace (handles tabs, multiple spaces, newlines) re.split(r"\s+", "hello world\tfoo\nbar") # ['hello', 'world', 'foo', 'bar'] # Strip multiple spaces down to one re.sub(r"\s+", " ", "hello world ").strip() # 'hello world' # Find URLs (a basic, common-case pattern — not RFC-complete) URL_RE = re.compile(r"https?://[^\s<>\"']+") URL_RE.findall("see https://example.com/x and http://foo.bar/baz here") # ['https://example.com/x', 'http://foo.bar/baz'] # Extract YYYY-MM-DD dates re.findall(r"\b(\d{4})-(\d{2})-(\d{2})\b", "log 2026-03-14 and 2025-12-01") # [('2026', '03', '14'), ('2025', '12', '01')] # Escape user input for safe inclusion in a pattern user_query = "C++ 3.0 (final)" safe = re.escape(user_query) # 'C\\+\\+\\ 3\\.0\\ \\(final\\)' # Now safe to embed: re.search(rf"^{safe}$", line)
re.escape() is essential when you splice user-supplied text into a pattern. Without it, a user typing . or ( would have those treated as regex metacharacters, and a user typing .* could match anything.
13. Common Mistakes
1. Forgetting the r prefix"\d" is a deprecated escape in regular strings — r"\d" is clean. The fix is one character; the bug it prevents can be subtle ("\b" becomes ASCII backspace).
2. Greedy quantifiers eating too much<.*> over <b>bold</b> matches the whole thing, not just <b>. Use <.*?> for lazy, or be explicit: <[^>]*>.
3. Anchoring without re.MULTILINE^ERROR on a multi-line log only matches the very first line. Pass re.MULTILINE when you mean per-line.
4. Unescaped special characters
Literal dot, plus, paren, brace, bracket, pipe — they all need \ to be themselves. Or use [.], [+], etc. Or re.escape(s) for runtime strings.
5. Treating regex as a parser
Repeat after me: HTML, JSON, CSV, source code, math expressions, balanced parens — not regex. Use the dedicated parser. Your team will thank you.
6. re.match vs re.search confusionmatch is anchored at the start of the string. If your pattern doesn't start with ^, match still requires the match to begin at index 0. Use search when you mean "find anywhere".
🎯 Your Turn — Find All Emails
Write find_emails(text) that returns a list of deduplicated, lowercased email addresses found in text. The regex should handle:
- Standard local parts:
alice@example.com - Plus-tags:
alice+filter@example.com - Dots in the local part:
first.last@example.com - Subdomains:
alice@mail.eng.example.com - Mixed case (should be returned lowercase)
Provide the regex inline with comments using re.VERBOSE.
text = """ Contact alice@example.com or ALICE@EXAMPLE.COM (same person). CC bob+team@mail.example.co.uk and first.last@sub.eng.example.com. Not an email: alice@, @example.com, just-a-word. """ # find_emails(text) → ['alice@example.com', # 'bob+team@mail.example.co.uk', # 'first.last@sub.eng.example.com']
Skeleton:
import re EMAIL_RE = re.compile(r""" # TODO 1: match the local part — letters/digits/dots/pluses/hyphens/underscores # TODO 2: an @ sign # TODO 3: a domain — one or more labels separated by dots # TODO 4: ensure a TLD of at least 2 letters at the end """, re.VERBOSE) def find_emails(text): # TODO 5: findall, lowercase, dedupe (preserve order) ...
Hint 1 — Local part character class
Allowed characters in the local part (for our purposes):[\w.+\-]+. \w covers letters, digits, and underscore; we add ., +, and - explicitly. Note: - inside a character class is fine at the end or escaped.
Hint 2 — Domain with multiple labels
A domain likemail.eng.example.com is a series of label-then-dot, ending in a final TLD label. One way: (?:[\w-]+\.)+[A-Za-z]{2,} — one or more "label." groups followed by a final 2+ letter TLD.
Hint 3 — Dedupe preserving order
Aset dedupes but loses order. To dedupe while preserving first-seen order: build a dict.fromkeys(items) and take its keys. Dicts in Python 3.7+ preserve insertion order.
Show full solution
import re EMAIL_RE = re.compile(r""" \b # word boundary so we don't match mid-word ( [\w.+\-]+ # local part: letters/digits/_/./+/- @ (?:[\w\-]+\.)+ # one or more "subdomain." labels [A-Za-z]{2,} # final TLD, 2+ letters (no digits) ) \b """, re.VERBOSE) def find_emails(text): """Return deduplicated, lowercase emails in first-seen order.""" matches = (m.group(1).lower() for m in EMAIL_RE.finditer(text)) return list(dict.fromkeys(matches)) # Demo text = """ Contact alice@example.com or ALICE@EXAMPLE.COM (same person). CC bob+team@mail.example.co.uk and first.last@sub.eng.example.com. Not an email: alice@, @example.com, just-a-word. """ for e in find_emails(text): print(e) # alice@example.com # bob+team@mail.example.co.uk # first.last@sub.eng.example.com
What's going on:
re.VERBOSElets us comment each piece of the pattern. This regex is now maintainable — six months from now you can still read it.\bword boundaries keep us from matching partial-word junk like the trailing punctuation in(alice@example.com).[\w.+\-]+is the local part. The\-is escaped (or could go last —[\w.+-]+works too); the others are literal inside the class.(?:[\w\-]+\.)+is "one or more label-then-dot" — a non-capturing group repeated. This handles arbitrary subdomain depth.[A-Za-z]{2,}anchors the TLD to be at least two letters — rules out junk likealice@example.1from passing.dict.fromkeys(...)is the canonical "ordered set" idiom in modern Python. We could've used a manual loop with aseenset, but this is one line.
Caveat: this regex is good for common cases, not RFC-compliant. RFC 5322 allows quoted local parts ("weird name"@example.com), IP literals (bob@[192.168.1.1]), and other oddities you'll almost never see in practice. If you need to "validate" an email for real, send a confirmation email — that's the only validator that's actually correct.
What You Learned
remodule with five functions:search,match,findall,finditer,sub.- Character classes (
\d \w \s .,[abc],[^abc],[a-z]) describe what one character can be. - Quantifiers (
? * + {n,m}) say how many; add?for lazy. - Anchors (
^ $ \b) match positions, not characters.\bis the underused gem. - Groups:
(...)captures,(?:...)doesn't,(?P<name>...)names. Backreference with\1or\g<name>. - Flags:
IGNORECASE,MULTILINE,DOTALL,VERBOSE. The last one makes complex patterns readable — use it. re.compilefor hot loops and for giving patterns first-class names.- Always
r"..."raw strings. Always. - Regex is for patterns, not structure. HTML, JSON, CSV, code — use real parsers.
Next: Datetime, Timezones, and Friends — handling dates and times properly, which is harder than regex by a margin you don't yet appreciate.
Practice this
on practicepython.inShort exercises that run in your browser and tell you what your code actually did, not just whether a test passed.