PythonMastery
intermediate 20 min read · lesson 10 of 13 in Python Intermediate

Regular Expressions

1 · The lesson

read

A regex is a tiny pattern language for describing shapes of text. "An email address." "A date in YYYY-MM-DD form." "Three digits followed by a dash." You write the shape; the engine finds (or rejects) matches.

Regex has a reputation for being write-only — and it earns it when people reach for it where they shouldn't. This lesson teaches the parts you'll actually use, when to use them, and — equally important — when to put the regex down and pick up a real parser.


1. The re Module — Your First Pattern

python
import re

text = "Order 4271 shipped on 2026-03-14"

m = re.search(r"\d+", text)            # find the first run of digits
print(m)                                # <re.Match object; span=(6, 10), match='4271'>
print(m.group())                        # '4271'
print(m.span())                         # (6, 10) — start and end indices

Three things to register from this example:

1. The r"..." raw string prefix. Always. Regex uses \ for everything (\d, \b, \w); without r, Python's string parser eats them first and you get bugs or warnings. We'll come back to this in Section 10.
2. re.search finds the first match anywhere in the string. It returns a Match object or None.
3. Match.group() gives you the matched text. Match.span() gives you the position.


2. The Five Functions You'll Use 95% of the Time

python
import re

text = "alice@example.com, bob@example.com, carol@other.org"

# 1. search — first match anywhere, or None
m = re.search(r"\w+@\w+\.\w+", text)
print(m.group())                        # alice@example.com

# 2. match — match only at the START of the string
print(re.match(r"\w+@", text))          # None — text starts with "alice@example.com," but match anchors at index 0
print(re.match(r"alice", text))         # Match object

# 3. findall — every match, as a list of strings
print(re.findall(r"\w+@\w+\.\w+", text))
# ['alice@example.com', 'bob@example.com', 'carol@other.org']

# 4. finditer — every match, as Match objects (preferred when you need positions)
for m in re.finditer(r"\w+@\w+\.\w+", text):
    print(m.group(), "at", m.span())

# 5. sub — find and replace
print(re.sub(r"\w+@\w+\.\w+", "[EMAIL]", text))
# [EMAIL], [EMAIL], [EMAIL]

search vs match is the rookie pitfall: match is anchored at the start, search is anywhere. When in doubt, search. match is mainly useful when you're parsing a known format from position 0 (like a line in a structured log).

findall is convenient but gives you only strings — no position info, and if you use groups it returns tuples of groups (no full match). finditer is what you reach for once you outgrow findall.


3. Character Classes

The building blocks for "what kind of character can be here?"

PatternMeans
\dA digit. Equivalent to [0-9] for ASCII.
\DA non-digit.
\wA "word" character — letter, digit, or underscore.
\WA non-word character.
\sWhitespace — space, tab, newline.
\SNon-whitespace.
.Any character except newline (unless re.DOTALL is set).
[abc]Any one of a, b, or c.
[^abc]Any character that is NOT a, b, or c.
[a-z]Any lowercase letter. Ranges work in sets.
[A-Za-z0-9_]Hand-rolled \w. Useful when you want to be explicit.
python
re.findall(r"\d", "abc123")             # ['1', '2', '3']
re.findall(r"\d+", "abc123 def456")     # ['123', '456']
re.findall(r"[aeiou]", "regex lesson")  # ['e', 'e', 'e', 'o']
re.findall(r"[^aeiou\s]", "regex")      # ['r', 'g', 'x'] — consonants
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

Inside [], most metacharacters lose their special meaning — [.] is a literal dot, [+] is a literal plus. The exceptions are ^ (negation, only at start), - (range), ] (closes the class), and \ (escape).


4. Quantifiers — How Many?

PatternMeans
?0 or 1 (optional)
*0 or more
+1 or more
{n}Exactly n
{n,m}Between n and m
{n,}At least n

All of these are greedy by default — they grab as much as possible before backtracking. Add ? after them to make them lazy (also called non-greedy or reluctant).

python
text = "<b>bold</b> and <i>italic</i>"

re.findall(r"<.*>", text)               # ['<b>bold</b> and <i>italic</i>']  — greedy
re.findall(r"<.*?>", text)              # ['<b>', '</b>', '<i>', '</i>']    — lazy
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

The greedy version starts at the first <, then .* eats everything including the rest of the angle brackets, then backtracks until it finds the last >. The lazy version stops at the first > after each <. Both are valid; you almost always want lazy for tag-like delimiters.

(That said: if you're actually parsing HTML, see Section 11. Don't.)


5. Anchors — Where in the String?

PatternMeans
^Start of string (or start of line with re.MULTILINE)
$End of string (or end of line with re.MULTILINE)
\bWord boundary — the transition between \w and \W
\BNon-word-boundary

The most underused of the four is \b. It matches the position between a word character and a non-word character (or string edge), with zero width. Brilliant for "match the word cat but not catalogue":

python
text = "cat caterpillar scattered cats"

re.findall(r"cat",   text)              # ['cat', 'cat', 'cat', 'cat']    — too many
re.findall(r"\bcat\b", text)            # ['cat']                          — just the word
re.findall(r"\bcat",   text)            # ['cat', 'cat', 'cat']            — word-start cats
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

^ and $ by default match the start/end of the whole string. With re.MULTILINE, they match the start/end of each line:

python
log = "ERROR: x\nINFO: y\nERROR: z"

re.findall(r"^ERROR", log)                          # ['ERROR']         — only line 1's ERROR
re.findall(r"^ERROR", log, re.MULTILINE)            # ['ERROR', 'ERROR']
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

6. Groups — Capturing Pieces

Parentheses (...) create a capture group — a sub-match you can pull out by index.

python
m = re.search(r"(\d{4})-(\d{2})-(\d{2})", "Released 2026-03-14, today.")
print(m.group(0))                       # 2026-03-14   — the full match
print(m.group(1))                       # 2026         — first group
print(m.group(2))                       # 03           — second group
print(m.group(3))                       # 14           — third group
print(m.groups())                       # ('2026', '03', '14')
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

Non-capturing groups (?:...) group without capturing — useful when you need grouping for alternation or quantifiers but don't want the result in .groups():

python
re.findall(r"(?:Mr|Mrs|Ms)\. (\w+)", "Mr. Stark, Ms. Potts")
# ['Stark', 'Potts']   — only the name is captured; the title isn't
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

Named groups (?P<name>...) are the readable upgrade. You access them by name instead of fragile indices:

python
m = re.search(r"(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})", "2026-03-14")
print(m.group("year"))                  # 2026
print(m.groupdict())                    # {'year': '2026', 'month': '03', 'day': '14'}
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

For anything more than two groups, name them. Index 1 vs index 2 confusion has burned everyone.


7. Backreferences

Once you've captured a group, you can refer to it later — both inside the pattern (to require a repeated match) and in a replacement string (to reuse the captured text).

python
# In the pattern: find words that are immediately repeated
re.findall(r"\b(\w+) \1\b", "the the cat sat on on the mat")
# ['the', 'on']

# In a replacement: swap first and last name
re.sub(r"(\w+) (\w+)", r"\2 \1", "Linus Teja")
# 'Teja Linus'

# Named variant
re.sub(r"(?P<first>\w+) (?P<last>\w+)",
       r"\g<last> \g<first>",
       "Linus Teja")
# 'Teja Linus'
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

\1 is the first captured group; \g<name> is the named-group form in replacement strings. Backreferences are how regex stops being purely "shape-matching" and starts being able to express relationships between parts of the input.


8. Flags

Flags modify how the engine interprets the pattern. Pass them as the third argument, or combine with |.

FlagEffect
re.IGNORECASE (or re.I)Case-insensitive matching.
re.MULTILINE (or re.M)^ and $ match line boundaries, not just string.
re.DOTALL (or re.S). matches newlines too.
re.VERBOSE (or re.X)Allow whitespace + # comments inside the pattern.
python
re.findall(r"hello", "Hello world", re.IGNORECASE)              # ['Hello']
re.findall(r".+", "line1\nline2", re.DOTALL)                    # ['line1\nline2']
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

re.VERBOSE is the one most people miss, and it turns intimidating patterns into readable ones:

python
phone_pattern = re.compile(r"""
    \+?                   # optional leading +
    (?P<country>\d{1,3})  # 1-3 digit country code
    [\s\-]?               # optional separator
    (?P<area>\d{3})       # area code
    [\s\-]?               # optional separator
    (?P<line>\d{4})       # line number
    """, re.VERBOSE)

m = phone_pattern.search("Call me at +1 555-7890 maybe")
print(m.groupdict())                    # {'country': '1', 'area': '555', 'line': '7890'}
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

Whitespace and #-comments inside the pattern are ignored. Want to match a literal space? \ or [ ]. Any non-trivial pattern that lives in your codebase for more than a week should be in re.VERBOSE form — your future self will read it instead of squinting at it.


9. re.compile — Precompile for Hot Loops

re.search, re.findall, etc. cache the compiled pattern internally — so calling them in a loop is fine, up to a point. For very hot loops, or to give a pattern a meaningful name, compile it once and reuse:

python
EMAIL_RE = re.compile(r"\b[\w.+-]+@[\w-]+\.[\w.-]+\b")

for line in big_log_file:
    for match in EMAIL_RE.finditer(line):
        process(match.group())
+ setup added so this can run · defines big_log_file, re, process
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

big_log_file = ["alpha", "beta", "gamma"]
re = _AutoMock('re')
def process(*_a, **_kw):
    print('-> process() called')
    return _AutoMock('process()')

Compiled patterns expose the same operations as module-level methods — EMAIL_RE.search, .findall, .finditer, .sub. They also make patterns first-class: you can pass them around, store them in dicts, and most importantly name them so the reader of your code knows what you're matching, not just how.


10. Always Use Raw Strings — r"..."

Without r, Python's string literal parser interprets backslashes first:

python
# WRONG — Python sees "\d" inside a normal string. "\d" isn't a known escape,
# so it stays as backslash-d, BUT this raises DeprecationWarning and will
# eventually be a SyntaxError. And for things like "\b" — that IS a known
# escape (backspace, \x08) — Python eats it before regex ever sees it.
pattern = "\bword\b"                    # backspace-word-backspace, not "\b word \b"

# RIGHT — raw string passes the backslashes through untouched
pattern = r"\bword\b"

The fix is one character. Add r in front of every regex pattern, every time. Make it muscle memory.

The same applies to replacement strings in re.sub if they contain backslashes:

python
re.sub(r"(\d+)", r"<\1>", "got 42 items")     # 'got <42> items'
+ setup added so this can run · defines re
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

re = _AutoMock('re')

11. When NOT to Use Regex

This section saves more pain than the previous ten combined. Regex matches patterns. It does not parse structure.

You want to…Use…Not regex because…
Parse HTML / XMLBeautifulSoup, lxml (see scraping)HTML is nested, regex isn't. You'll get the simple case and the hard case will eat your weekend.
Parse JSONjson moduleJSON is structured; json.loads validates and decodes in one call.
Parse CSVcsv moduleEmbedded commas, quoted fields, multi-line cells — csv handles it; regex will miss edge cases.
Validate emails "properly"Send a confirmation emailThe RFC 5322 grammar is ~6 pages. Any regex you find online is wrong in some case. The only true validator is "does mail get delivered?"
Parse Python / SQL / any codeAn actual parser (ast, sqlglot)Code is hierarchical; regex is flat.
Extract a number from a sentenceRegex ✓This is genuinely a flat pattern. Go for it.

The rule of thumb: if the input has nested or hierarchical structure, regex is the wrong tool. If it has flat, locally-recognisable patterns, regex is excellent.


12. Real-World Patterns You'll Actually Reuse

python
import re

# Split on any run of whitespace (handles tabs, multiple spaces, newlines)
re.split(r"\s+", "hello   world\tfoo\nbar")
# ['hello', 'world', 'foo', 'bar']

# Strip multiple spaces down to one
re.sub(r"\s+", " ", "hello    world  ").strip()
# 'hello world'

# Find URLs (a basic, common-case pattern — not RFC-complete)
URL_RE = re.compile(r"https?://[^\s<>\"']+")
URL_RE.findall("see https://example.com/x and http://foo.bar/baz here")
# ['https://example.com/x', 'http://foo.bar/baz']

# Extract YYYY-MM-DD dates
re.findall(r"\b(\d{4})-(\d{2})-(\d{2})\b", "log 2026-03-14 and 2025-12-01")
# [('2026', '03', '14'), ('2025', '12', '01')]

# Escape user input for safe inclusion in a pattern
user_query = "C++ 3.0 (final)"
safe = re.escape(user_query)
# 'C\\+\\+\\ 3\\.0\\ \\(final\\)'
# Now safe to embed: re.search(rf"^{safe}$", line)

re.escape() is essential when you splice user-supplied text into a pattern. Without it, a user typing . or ( would have those treated as regex metacharacters, and a user typing .* could match anything.


13. Common Mistakes

1. Forgetting the r prefix
"\d" is a deprecated escape in regular strings — r"\d" is clean. The fix is one character; the bug it prevents can be subtle ("\b" becomes ASCII backspace).

2. Greedy quantifiers eating too much
<.*> over <b>bold</b> matches the whole thing, not just <b>. Use <.*?> for lazy, or be explicit: <[^>]*>.

3. Anchoring without re.MULTILINE
^ERROR on a multi-line log only matches the very first line. Pass re.MULTILINE when you mean per-line.

4. Unescaped special characters
Literal dot, plus, paren, brace, bracket, pipe — they all need \ to be themselves. Or use [.], [+], etc. Or re.escape(s) for runtime strings.

5. Treating regex as a parser
Repeat after me: HTML, JSON, CSV, source code, math expressions, balanced parens — not regex. Use the dedicated parser. Your team will thank you.

6. re.match vs re.search confusion
match is anchored at the start of the string. If your pattern doesn't start with ^, match still requires the match to begin at index 0. Use search when you mean "find anywhere".


🎯 Your Turn — Find All Emails

Write find_emails(text) that returns a list of deduplicated, lowercased email addresses found in text. The regex should handle:

  • Standard local parts: alice@example.com
  • Plus-tags: alice+filter@example.com
  • Dots in the local part: first.last@example.com
  • Subdomains: alice@mail.eng.example.com
  • Mixed case (should be returned lowercase)

Provide the regex inline with comments using re.VERBOSE.

python
text = """
Contact alice@example.com or ALICE@EXAMPLE.COM (same person).
CC bob+team@mail.example.co.uk and first.last@sub.eng.example.com.
Not an email: alice@, @example.com, just-a-word.
"""
# find_emails(text) → ['alice@example.com',
#                      'bob+team@mail.example.co.uk',
#                      'first.last@sub.eng.example.com']

Skeleton:

python
import re

EMAIL_RE = re.compile(r"""
    # TODO 1: match the local part — letters/digits/dots/pluses/hyphens/underscores
    # TODO 2: an @ sign
    # TODO 3: a domain — one or more labels separated by dots
    # TODO 4: ensure a TLD of at least 2 letters at the end
    """, re.VERBOSE)

def find_emails(text):
    # TODO 5: findall, lowercase, dedupe (preserve order)
    ...
Hint 1 — Local part character class Allowed characters in the local part (for our purposes): [\w.+\-]+. \w covers letters, digits, and underscore; we add ., +, and - explicitly. Note: - inside a character class is fine at the end or escaped.
Hint 2 — Domain with multiple labels A domain like mail.eng.example.com is a series of label-then-dot, ending in a final TLD label. One way: (?:[\w-]+\.)+[A-Za-z]{2,} — one or more "label." groups followed by a final 2+ letter TLD.
Hint 3 — Dedupe preserving order A set dedupes but loses order. To dedupe while preserving first-seen order: build a dict.fromkeys(items) and take its keys. Dicts in Python 3.7+ preserve insertion order.
Show full solution
python
import re

EMAIL_RE = re.compile(r"""
    \b                          # word boundary so we don't match mid-word
    (
        [\w.+\-]+               # local part: letters/digits/_/./+/-
        @
        (?:[\w\-]+\.)+          # one or more "subdomain." labels
        [A-Za-z]{2,}            # final TLD, 2+ letters (no digits)
    )
    \b
    """, re.VERBOSE)

def find_emails(text):
    """Return deduplicated, lowercase emails in first-seen order."""
    matches = (m.group(1).lower() for m in EMAIL_RE.finditer(text))
    return list(dict.fromkeys(matches))


# Demo
text = """
Contact alice@example.com or ALICE@EXAMPLE.COM (same person).
CC bob+team@mail.example.co.uk and first.last@sub.eng.example.com.
Not an email: alice@, @example.com, just-a-word.
"""

for e in find_emails(text):
    print(e)
# alice@example.com
# bob+team@mail.example.co.uk
# first.last@sub.eng.example.com

What's going on:

  • re.VERBOSE lets us comment each piece of the pattern. This regex is now maintainable — six months from now you can still read it.
  • \b word boundaries keep us from matching partial-word junk like the trailing punctuation in (alice@example.com).
  • [\w.+\-]+ is the local part. The \- is escaped (or could go last — [\w.+-]+ works too); the others are literal inside the class.
  • (?:[\w\-]+\.)+ is "one or more label-then-dot" — a non-capturing group repeated. This handles arbitrary subdomain depth.
  • [A-Za-z]{2,} anchors the TLD to be at least two letters — rules out junk like alice@example.1 from passing.
  • dict.fromkeys(...) is the canonical "ordered set" idiom in modern Python. We could've used a manual loop with a seen set, but this is one line.

Caveat: this regex is good for common cases, not RFC-compliant. RFC 5322 allows quoted local parts ("weird name"@example.com), IP literals (bob@[192.168.1.1]), and other oddities you'll almost never see in practice. If you need to "validate" an email for real, send a confirmation email — that's the only validator that's actually correct.


What You Learned

  • re module with five functions: search, match, findall, finditer, sub.
  • Character classes (\d \w \s ., [abc], [^abc], [a-z]) describe what one character can be.
  • Quantifiers (? * + {n,m}) say how many; add ? for lazy.
  • Anchors (^ $ \b) match positions, not characters. \b is the underused gem.
  • Groups: (...) captures, (?:...) doesn't, (?P<name>...) names. Backreference with \1 or \g<name>.
  • Flags: IGNORECASE, MULTILINE, DOTALL, VERBOSE. The last one makes complex patterns readable — use it.
  • re.compile for hot loops and for giving patterns first-class names.
  • Always r"..." raw strings. Always.
  • Regex is for patterns, not structure. HTML, JSON, CSV, code — use real parsers.

Next: Datetime, Timezones, and Friends — handling dates and times properly, which is harder than regex by a margin you don't yet appreciate.

Practice this

on practicepython.in

Short exercises that run in your browser and tell you what your code actually did, not just whether a test passed.