PythonMastery
intermediate 30 min read · lesson 13 of 15 in Projects

Project: Markdown to HTML Converter

1 · The lesson

read

You'll hand-roll a Markdown-to-HTML converter — no markdown library, no mistune. Headers, bold/italic, lists, code blocks, links, images. By the end you'll have written a small interpreter, learned the state-machine pattern for line-by-line parsers, and handled an HTML-escaping subtlety that real libraries occasionally get wrong.

What you'll practice: regex pipelines, line-by-line state machines, the small-interpreter pattern, HTML escaping, class-based stateful processing.


Step 1 — Inline Replacements

Start with the easy stuff: **bold**, *italic*, `code`. All regex.

python
import re

def render_inline(text):
    # Order matters: bold (**) before italic (*) so we don't half-match
    text = re.sub(r"\*\*(.+?)\*\*", r"<strong>\1</strong>", text)
    text = re.sub(r"\*(.+?)\*",     r"<em>\1</em>",         text)
    text = re.sub(r"`(.+?)`",       r"<code>\1</code>",     text)
    return text

# Test
print(render_inline("This is **bold** and *italic* and `code`."))
print(render_inline("Mix **bold *italic* stuff**."))

The .+? is non-greedy — match as few chars as possible. Without it, **a** and **b** would match the whole thing as one <strong> block. Non-greedy is essential whenever delimiters can appear in pairs.


Step 2 — Headers

# Header → <h1>Header</h1>, six levels deep. Line-by-line, one regex per pass.

python
import re

def render_header(line):
    m = re.match(r"^(#{1,6})\s+(.+?)\s*$", line)
    if not m:
        return None        # not a header
    level = len(m.group(1))
    return f"<h{level}>{m.group(2)}</h{level}>"

# Test
for line in [
    "# Top-level title",
    "## Section",
    "###### Tiny header",
    "####### Too many hashes",       # 7 hashes — not a header
    "not a header",
]:
    out = render_header(line)
    print(f"  {line!r:<35} → {out!r}")

The regex breaks down as:


  • ^(#{1,6}) — start of line, 1-6 hashes captured

  • \s+ — at least one space (so #foo is NOT a header)

  • (.+?)\s*$ — the title, trimmed of trailing whitespace

Returning None for "not a header" lets the caller fall through to the next rule. Classic dispatcher shape.


Step 3 — Lists (Where State Machines Earn Their Keep)

Markdown's tricky bit: a list is a block, not a per-line thing. <ul> opens before the first item, closes after the last. To detect "the last", you need state.

python
import re

def render(lines):
    """Convert lines of markdown to lines of HTML, handling lists."""
    out = []
    in_list = None                # None, 'ul', or 'ol'

    for line in lines:
        # Bullet list item?
        m_ul = re.match(r"^\s*[-*]\s+(.+)$", line)
        # Numbered list item?
        m_ol = re.match(r"^\s*\d+\.\s+(.+)$", line)

        if m_ul:
            if in_list != "ul":
                if in_list: out.append(f"</{in_list}>")
                out.append("<ul>")
                in_list = "ul"
            out.append(f"  <li>{m_ul.group(1)}</li>")
        elif m_ol:
            if in_list != "ol":
                if in_list: out.append(f"</{in_list}>")
                out.append("<ol>")
                in_list = "ol"
            out.append(f"  <li>{m_ol.group(1)}</li>")
        else:
            if in_list:
                out.append(f"</{in_list}>")
                in_list = None
            out.append(line)

    # End-of-document: don't forget to close any open list
    if in_list:
        out.append(f"</{in_list}>")
    return "\n".join(out)

# Test
md = [
    "Some intro text.",
    "- apples",
    "- bananas",
    "- cherries",
    "",
    "1. first",
    "2. second",
    "",
    "Some closing text.",
]
print(render(md))

in_list is the state variable. Three states (None, "ul", "ol") with transitions on each input line. This shape generalises to any block-level parser: an outer state, transitions triggered by pattern matches, and a "don't forget the end" cleanup pass.


Step 4 — Code Blocks (Another State Machine)

Fenced code blocks (```...```) are another stateful pattern — content inside the fence is verbatim, NOT processed for Markdown.

python
def render_with_code(lines):
    out = []
    in_code = False

    for line in lines:
        if line.strip().startswith("```"):
            if not in_code:
                out.append("<pre><code>")
                in_code = True
            else:
                out.append("</code></pre>")
                in_code = False
            continue

        if in_code:
            # Verbatim — no markdown processing inside code blocks
            out.append(line)
        else:
            out.append(line)        # would run inline/header rules here

    if in_code:
        out.append("</code></pre>")  # auto-close unterminated block
    return "\n".join(out)

md = [
    "Here's some code:",
    "```",
    "def greet(name):",
    "    return f'hello, {name}'",
    "```",
    "And we're back to prose with **bold**.",
]
print(render_with_code(md))

The toggle pattern (in_code = not in_code would work too) is the simplest possible state machine — two states, one transition. But the principle is identical to ten-state parsers: "what state am I in, and what does this input do to that state?"


Inline patterns again — but with capture groups for the parts.

python
import re

def render_links(text):
    # Images FIRST — they start with !, share the syntax otherwise
    text = re.sub(
        r"!\[([^\]]*)\]\(([^)]+)\)",
        r'<img alt="\1" src="\2">',
        text,
    )
    # Then links
    text = re.sub(
        r"\[([^\]]+)\]\(([^)]+)\)",
        r'<a href="\2">\1</a>',
        text,
    )
    return text

# Test
samples = [
    "See [the docs](https://docs.python.org).",
    "An image: ![diagram](/img/arch.png) inline.",
    "Mixed [link](http://x.com) and ![pic](pic.jpg).",
]
for s in samples:
    print(f"  {s}")
    print(f"  → {render_links(s)}\n")

Order matters again: process images before links, otherwise the leading ! confuses the link regex.

[^\]]+ means "one or more chars that aren't ]" — so the regex doesn't run past the closing bracket of the label. The same trick ([^)]+ for the URL) keeps the URL match tight.


Step 6 — Polished Final Version (a Markdown Class)

Put it all together: paragraphs, all the state machines coordinated, HTML escaping in prose but NOT inside code blocks.

python
import re
from html import escape

class Markdown:
    """Tiny subset-of-CommonMark Markdown → HTML converter.

    Handles: headers, paragraphs, bold/italic/code, lists, code blocks,
    links and images. HTML chars are escaped EXCEPT inside fenced code.
    """

    HEADER_RE = re.compile(r"^(#{1,6})\s+(.+?)\s*$")
    UL_RE     = re.compile(r"^\s*[-*]\s+(.+)$")
    OL_RE     = re.compile(r"^\s*\d+\.\s+(.+)$")
    FENCE_RE  = re.compile(r"^\s*```")

    def to_html(self, text):
        lines = text.splitlines()
        out = []
        self._in_list = None
        self._in_code = False
        self._para = []

        for line in lines:
            if self.FENCE_RE.match(line):
                self._flush_para(out)
                self._close_list(out)
                if not self._in_code:
                    out.append("<pre><code>")
                    self._in_code = True
                else:
                    out.append("</code></pre>")
                    self._in_code = False
                continue

            if self._in_code:
                # Escape inside code (avoids breaking <pre>), but no markdown.
                out.append(escape(line))
                continue

            if not line.strip():
                self._flush_para(out)
                self._close_list(out)
                continue

            m_h = self.HEADER_RE.match(line)
            if m_h:
                self._flush_para(out)
                self._close_list(out)
                level = len(m_h.group(1))
                out.append(f"<h{level}>{self._inline(m_h.group(2))}</h{level}>")
                continue

            m_ul = self.UL_RE.match(line)
            m_ol = self.OL_RE.match(line)
            if m_ul or m_ol:
                self._flush_para(out)
                kind = "ul" if m_ul else "ol"
                if self._in_list != kind:
                    self._close_list(out)
                    out.append(f"<{kind}>")
                    self._in_list = kind
                text = (m_ul or m_ol).group(1)
                out.append(f"  <li>{self._inline(text)}</li>")
                continue

            # Default: paragraph text. Accumulate until blank line.
            self._close_list(out)
            self._para.append(line)

        # End of document — flush anything still open
        self._flush_para(out)
        self._close_list(out)
        if self._in_code:
            out.append("</code></pre>")

        return "\n".join(out)

    def _inline(self, text):
        # Escape FIRST, then re-introduce our HTML via the replacements.
        # (Escaping after would mangle <strong> etc.)
        text = escape(text)
        text = re.sub(r"!\[([^\]]*)\]\(([^)]+)\)",
                      r'<img alt="\1" src="\2">', text)
        text = re.sub(r"\[([^\]]+)\]\(([^)]+)\)",
                      r'<a href="\2">\1</a>', text)
        text = re.sub(r"\*\*(.+?)\*\*", r"<strong>\1</strong>", text)
        text = re.sub(r"\*(.+?)\*",     r"<em>\1</em>",         text)
        text = re.sub(r"`(.+?)`",       r"<code>\1</code>",     text)
        return text

    def _flush_para(self, out):
        if self._para:
            body = " ".join(self._para)
            out.append(f"<p>{self._inline(body)}</p>")
            self._para = []

    def _close_list(self, out):
        if self._in_list:
            out.append(f"</{self._in_list}>")
            self._in_list = None


# Demo
SAMPLE = """# Hello, Markdown

This is a **bold** intro with *italic* and `inline code`. We also
support [links](https://example.com) and ![images](pic.png).

## Lists

- one
- two
- three

1. first
2. second

## Code
def safe(): return ""
python
The line above is escaped. & < > all become entities.
"""

print(Markdown().to_html(SAMPLE))

The security insight: escape() runs on prose so that <script>alert(1)</script> in user input becomes &lt;script&gt;... — safe to embed. But it ALSO runs on the body of code blocks, since otherwise the literal < in code would break the surrounding <pre>. The bug to avoid: many home-grown parsers escape the prose but forget the code blocks, so a < in a code sample silently breaks the HTML downstream. The reverse bug — escaping the prose AFTER converting ** to <strong> — destroys the tags entirely.


Stretch Goals

1. GitHub-flavoured tables: | col1 | col2 |\n|------|------|\n| a | b | → <table>. Another state machine: detect the header / separator / row sequence.
2. Task lists: - [ ] todo and - [x] done → <input type="checkbox" disabled> inside the <li>. Pre-process the list-item text.
3. Syntax highlighting: hand off code blocks to Pygments (pip install pygments). Wrap the fence in <pre class="highlight"><code class="language-python">...</code></pre>.
4. Footnote support: text[^1] ... [^1]: explanation → numbered footnotes at the bottom of the document. Two-pass: collect definitions, then render.
5. Stream a huge file: process line by line with a generator, yielding HTML lines as you go. Lets you convert 100 MB Markdown without loading it all into memory.


🎯 Your Turn — slugify_headers(html)

GitHub READMEs let you deep-link to a header: #installation jumps to the <h2>Installation</h2>. To enable that, the converter has to add id="installation" to each header. Build a post-processor that does exactly this.

python
import re

def slugify_headers(html):
    """Add id='...' to every <h1>-<h6> tag based on its text content.

    Rules for the slug:
      - lowercase the header text
      - replace runs of whitespace with a single hyphen
      - strip any chars that aren't letters, digits, or hyphens

    Example:
      <h2>Getting Started!</h2>  →  <h2 id="getting-started">Getting Started!</h2>
      <h1>Step 3 — Setup</h1>    →  <h1 id="step-3-setup">Step 3 — Setup</h1>
    """
    # TODO 1: write a helper slugify(s) that applies the three rules
    # TODO 2: use re.sub with a CALLBACK function to find each <hN>TEXT</hN>
    #         The callback receives the match, computes the slug from
    #         match.group(2), and returns the replacement string.
    # TODO 3: think carefully about the pattern:
    #           <(h[1-6])>([^<]+)</\1>
    pass


# Test
html = """<h1>Project Title</h1>
<p>Some intro.</p>
<h2>Step 1 — Setup</h2>
<p>...</p>
<h2>Step 2: The Plan</h2>
<h3>A Sub-section!</h3>"""

print(slugify_headers(html))
Hint 1 — Building the slug s = s.lower(); s = re.sub(r"\s+", "-", s); s = re.sub(r"[^a-z0-9-]", "", s). The order matters — lowercase before the alphanumeric filter, replace whitespace before stripping (otherwise multi-word headers collapse).
Hint 2 — Regex with a callback re.sub accepts a function as the replacement. The function receives the Match object and returns the replacement string: re.sub(r"<(h[1-6])>([^<]+)</\1>", lambda m: f'<{m[1]} id="{slugify(m[2])}">{m[2]}</{m[1]}>', html). The \1 backreference in the pattern ensures the opening and closing tags match (you don't slugify <h1>...</h2>).
Show full solution
python
import re

def slugify(s):
    s = s.lower()
    s = re.sub(r"\s+", "-", s)
    s = re.sub(r"[^a-z0-9-]", "", s)
    return s

def slugify_headers(html):
    def replace(m):
        tag, text = m.group(1), m.group(2)
        return f'<{tag} id="{slugify(text)}">{text}</{tag}>'
    return re.sub(r"<(h[1-6])>([^<]+)</\1>", replace, html)


html = """<h1>Project Title</h1>
<p>Some intro.</p>
<h2>Step 1 — Setup</h2>
<p>...</p>
<h2>Step 2: The Plan</h2>
<h3>A Sub-section!</h3>"""

print(slugify_headers(html))

Output:

python
<h1 id="project-title">Project Title</h1>
<p>Some intro.</p>
<h2 id="step-1--setup">Step 1 — Setup</h2>
<p>...</p>
<h2 id="step-2-the-plan">Step 2: The Plan</h2>
<h3 id="a-sub-section">A Sub-section!</h3>

Note the double-hyphen in step-1--setup — the em-dash got stripped but the spaces around it became hyphens. Real slugifiers collapse runs of hyphens too: s = re.sub(r"-+", "-", s). Add that as a polish step if you want exact GitHub parity.

The callback-replacement pattern is the single most useful re.sub feature. Any time the replacement depends on the content of the match (not just rearranging capture groups), reach for it.


What You Learned

  • Inline regex pipelines — non-greedy .+?, ordered passes (bold before italic, images before links).
  • The state-machine pattern for block-level parsers: state variable, transitions per input line, end-of-document cleanup.
  • HTML escaping discipline — escape prose, escape code-block bodies, NEVER escape after generating tags.
  • re.sub with a callback — when the replacement depends on the match content.
  • The class-as-state-bag pattern — self._in_list, self._in_code, self._para instead of dragging four variables through every helper.

You just built a tiny interpreter. It tokenises (regex), keeps state (the flags and the paragraph buffer), and emits output (the HTML lines). Compilers, template engines, and config parsers all share this shape — once you've written one, every other one feels familiar.

That's the Projects track for now. From here, the next levels build out web (Flask, FastAPI), data (pandas, numpy), and async topics — each project will use the foundations you've practiced across these 13 builds.

← Back to Academy

Practice this

on practicepython.in

Short exercises that run in your browser and tell you what your code actually did, not just whether a test passed.