Project: Markdown to HTML Converter
1 · The lesson
readYou'll hand-roll a Markdown-to-HTML converter — no markdown library, no mistune. Headers, bold/italic, lists, code blocks, links, images. By the end you'll have written a small interpreter, learned the state-machine pattern for line-by-line parsers, and handled an HTML-escaping subtlety that real libraries occasionally get wrong.
What you'll practice: regex pipelines, line-by-line state machines, the small-interpreter pattern, HTML escaping, class-based stateful processing.
Step 1 — Inline Replacements
Start with the easy stuff: **bold**, *italic*, `code`. All regex.
import re def render_inline(text): # Order matters: bold (**) before italic (*) so we don't half-match text = re.sub(r"\*\*(.+?)\*\*", r"<strong>\1</strong>", text) text = re.sub(r"\*(.+?)\*", r"<em>\1</em>", text) text = re.sub(r"`(.+?)`", r"<code>\1</code>", text) return text # Test print(render_inline("This is **bold** and *italic* and `code`.")) print(render_inline("Mix **bold *italic* stuff**."))
The .+? is non-greedy — match as few chars as possible. Without it, **a** and **b** would match the whole thing as one <strong> block. Non-greedy is essential whenever delimiters can appear in pairs.
Step 2 — Headers
# Header → <h1>Header</h1>, six levels deep. Line-by-line, one regex per pass.
import re def render_header(line): m = re.match(r"^(#{1,6})\s+(.+?)\s*$", line) if not m: return None # not a header level = len(m.group(1)) return f"<h{level}>{m.group(2)}</h{level}>" # Test for line in [ "# Top-level title", "## Section", "###### Tiny header", "####### Too many hashes", # 7 hashes — not a header "not a header", ]: out = render_header(line) print(f" {line!r:<35} → {out!r}")
The regex breaks down as:
^(#{1,6})— start of line, 1-6 hashes captured\s+— at least one space (so#foois NOT a header)(.+?)\s*$— the title, trimmed of trailing whitespace
Returning None for "not a header" lets the caller fall through to the next rule. Classic dispatcher shape.
Step 3 — Lists (Where State Machines Earn Their Keep)
Markdown's tricky bit: a list is a block, not a per-line thing. <ul> opens before the first item, closes after the last. To detect "the last", you need state.
import re def render(lines): """Convert lines of markdown to lines of HTML, handling lists.""" out = [] in_list = None # None, 'ul', or 'ol' for line in lines: # Bullet list item? m_ul = re.match(r"^\s*[-*]\s+(.+)$", line) # Numbered list item? m_ol = re.match(r"^\s*\d+\.\s+(.+)$", line) if m_ul: if in_list != "ul": if in_list: out.append(f"</{in_list}>") out.append("<ul>") in_list = "ul" out.append(f" <li>{m_ul.group(1)}</li>") elif m_ol: if in_list != "ol": if in_list: out.append(f"</{in_list}>") out.append("<ol>") in_list = "ol" out.append(f" <li>{m_ol.group(1)}</li>") else: if in_list: out.append(f"</{in_list}>") in_list = None out.append(line) # End-of-document: don't forget to close any open list if in_list: out.append(f"</{in_list}>") return "\n".join(out) # Test md = [ "Some intro text.", "- apples", "- bananas", "- cherries", "", "1. first", "2. second", "", "Some closing text.", ] print(render(md))
in_list is the state variable. Three states (None, "ul", "ol") with transitions on each input line. This shape generalises to any block-level parser: an outer state, transitions triggered by pattern matches, and a "don't forget the end" cleanup pass.
Step 4 — Code Blocks (Another State Machine)
Fenced code blocks (```...```) are another stateful pattern — content inside the fence is verbatim, NOT processed for Markdown.
def render_with_code(lines): out = [] in_code = False for line in lines: if line.strip().startswith("```"): if not in_code: out.append("<pre><code>") in_code = True else: out.append("</code></pre>") in_code = False continue if in_code: # Verbatim — no markdown processing inside code blocks out.append(line) else: out.append(line) # would run inline/header rules here if in_code: out.append("</code></pre>") # auto-close unterminated block return "\n".join(out) md = [ "Here's some code:", "```", "def greet(name):", " return f'hello, {name}'", "```", "And we're back to prose with **bold**.", ] print(render_with_code(md))
The toggle pattern (in_code = not in_code would work too) is the simplest possible state machine — two states, one transition. But the principle is identical to ten-state parsers: "what state am I in, and what does this input do to that state?"
Step 5 — Links and Images
Inline patterns again — but with capture groups for the parts.
import re def render_links(text): # Images FIRST — they start with !, share the syntax otherwise text = re.sub( r"!\[([^\]]*)\]\(([^)]+)\)", r'<img alt="\1" src="\2">', text, ) # Then links text = re.sub( r"\[([^\]]+)\]\(([^)]+)\)", r'<a href="\2">\1</a>', text, ) return text # Test samples = [ "See [the docs](https://docs.python.org).", "An image:  inline.", "Mixed [link](http://x.com) and .", ] for s in samples: print(f" {s}") print(f" → {render_links(s)}\n")
Order matters again: process images before links, otherwise the leading ! confuses the link regex.
[^\]]+ means "one or more chars that aren't ]" — so the regex doesn't run past the closing bracket of the label. The same trick ([^)]+ for the URL) keeps the URL match tight.
Step 6 — Polished Final Version (a Markdown Class)
Put it all together: paragraphs, all the state machines coordinated, HTML escaping in prose but NOT inside code blocks.
import re from html import escape class Markdown: """Tiny subset-of-CommonMark Markdown → HTML converter. Handles: headers, paragraphs, bold/italic/code, lists, code blocks, links and images. HTML chars are escaped EXCEPT inside fenced code. """ HEADER_RE = re.compile(r"^(#{1,6})\s+(.+?)\s*$") UL_RE = re.compile(r"^\s*[-*]\s+(.+)$") OL_RE = re.compile(r"^\s*\d+\.\s+(.+)$") FENCE_RE = re.compile(r"^\s*```") def to_html(self, text): lines = text.splitlines() out = [] self._in_list = None self._in_code = False self._para = [] for line in lines: if self.FENCE_RE.match(line): self._flush_para(out) self._close_list(out) if not self._in_code: out.append("<pre><code>") self._in_code = True else: out.append("</code></pre>") self._in_code = False continue if self._in_code: # Escape inside code (avoids breaking <pre>), but no markdown. out.append(escape(line)) continue if not line.strip(): self._flush_para(out) self._close_list(out) continue m_h = self.HEADER_RE.match(line) if m_h: self._flush_para(out) self._close_list(out) level = len(m_h.group(1)) out.append(f"<h{level}>{self._inline(m_h.group(2))}</h{level}>") continue m_ul = self.UL_RE.match(line) m_ol = self.OL_RE.match(line) if m_ul or m_ol: self._flush_para(out) kind = "ul" if m_ul else "ol" if self._in_list != kind: self._close_list(out) out.append(f"<{kind}>") self._in_list = kind text = (m_ul or m_ol).group(1) out.append(f" <li>{self._inline(text)}</li>") continue # Default: paragraph text. Accumulate until blank line. self._close_list(out) self._para.append(line) # End of document — flush anything still open self._flush_para(out) self._close_list(out) if self._in_code: out.append("</code></pre>") return "\n".join(out) def _inline(self, text): # Escape FIRST, then re-introduce our HTML via the replacements. # (Escaping after would mangle <strong> etc.) text = escape(text) text = re.sub(r"!\[([^\]]*)\]\(([^)]+)\)", r'<img alt="\1" src="\2">', text) text = re.sub(r"\[([^\]]+)\]\(([^)]+)\)", r'<a href="\2">\1</a>', text) text = re.sub(r"\*\*(.+?)\*\*", r"<strong>\1</strong>", text) text = re.sub(r"\*(.+?)\*", r"<em>\1</em>", text) text = re.sub(r"`(.+?)`", r"<code>\1</code>", text) return text def _flush_para(self, out): if self._para: body = " ".join(self._para) out.append(f"<p>{self._inline(body)}</p>") self._para = [] def _close_list(self, out): if self._in_list: out.append(f"</{self._in_list}>") self._in_list = None # Demo SAMPLE = """# Hello, Markdown This is a **bold** intro with *italic* and `inline code`. We also support [links](https://example.com) and . ## Lists - one - two - three 1. first 2. second ## Code
The line above is escaped. & < > all become entities. """ print(Markdown().to_html(SAMPLE))
The security insight: escape() runs on prose so that <script>alert(1)</script> in user input becomes <script>... — safe to embed. But it ALSO runs on the body of code blocks, since otherwise the literal < in code would break the surrounding <pre>. The bug to avoid: many home-grown parsers escape the prose but forget the code blocks, so a < in a code sample silently breaks the HTML downstream. The reverse bug — escaping the prose AFTER converting ** to <strong> — destroys the tags entirely.
Stretch Goals
1. GitHub-flavoured tables: | col1 | col2 |\n|------|------|\n| a | b | → <table>. Another state machine: detect the header / separator / row sequence.
2. Task lists: - [ ] todo and - [x] done → <input type="checkbox" disabled> inside the <li>. Pre-process the list-item text.
3. Syntax highlighting: hand off code blocks to Pygments (pip install pygments). Wrap the fence in <pre class="highlight"><code class="language-python">...</code></pre>.
4. Footnote support: text[^1] ... [^1]: explanation → numbered footnotes at the bottom of the document. Two-pass: collect definitions, then render.
5. Stream a huge file: process line by line with a generator, yielding HTML lines as you go. Lets you convert 100 MB Markdown without loading it all into memory.
🎯 Your Turn — slugify_headers(html)
GitHub READMEs let you deep-link to a header: #installation jumps to the <h2>Installation</h2>. To enable that, the converter has to add id="installation" to each header. Build a post-processor that does exactly this.
import re def slugify_headers(html): """Add id='...' to every <h1>-<h6> tag based on its text content. Rules for the slug: - lowercase the header text - replace runs of whitespace with a single hyphen - strip any chars that aren't letters, digits, or hyphens Example: <h2>Getting Started!</h2> → <h2 id="getting-started">Getting Started!</h2> <h1>Step 3 — Setup</h1> → <h1 id="step-3-setup">Step 3 — Setup</h1> """ # TODO 1: write a helper slugify(s) that applies the three rules # TODO 2: use re.sub with a CALLBACK function to find each <hN>TEXT</hN> # The callback receives the match, computes the slug from # match.group(2), and returns the replacement string. # TODO 3: think carefully about the pattern: # <(h[1-6])>([^<]+)</\1> pass # Test html = """<h1>Project Title</h1> <p>Some intro.</p> <h2>Step 1 — Setup</h2> <p>...</p> <h2>Step 2: The Plan</h2> <h3>A Sub-section!</h3>""" print(slugify_headers(html))
Hint 1 — Building the slug
s = s.lower(); s = re.sub(r"\s+", "-", s); s = re.sub(r"[^a-z0-9-]", "", s).
The order matters — lowercase before the alphanumeric filter, replace whitespace before stripping (otherwise multi-word headers collapse).
Hint 2 — Regex with a callback
re.sub accepts a function as the replacement. The function receives the Match object and returns the replacement string:
re.sub(r"<(h[1-6])>([^<]+)</\1>", lambda m: f'<{m[1]} id="{slugify(m[2])}">{m[2]}</{m[1]}>', html).
The \1 backreference in the pattern ensures the opening and closing tags match (you don't slugify <h1>...</h2>).
Show full solution
import re def slugify(s): s = s.lower() s = re.sub(r"\s+", "-", s) s = re.sub(r"[^a-z0-9-]", "", s) return s def slugify_headers(html): def replace(m): tag, text = m.group(1), m.group(2) return f'<{tag} id="{slugify(text)}">{text}</{tag}>' return re.sub(r"<(h[1-6])>([^<]+)</\1>", replace, html) html = """<h1>Project Title</h1> <p>Some intro.</p> <h2>Step 1 — Setup</h2> <p>...</p> <h2>Step 2: The Plan</h2> <h3>A Sub-section!</h3>""" print(slugify_headers(html))
Output:
<h1 id="project-title">Project Title</h1> <p>Some intro.</p> <h2 id="step-1--setup">Step 1 — Setup</h2> <p>...</p> <h2 id="step-2-the-plan">Step 2: The Plan</h2> <h3 id="a-sub-section">A Sub-section!</h3>
Note the double-hyphen in step-1--setup — the em-dash got stripped but the spaces around it became hyphens. Real slugifiers collapse runs of hyphens too: s = re.sub(r"-+", "-", s). Add that as a polish step if you want exact GitHub parity.
The callback-replacement pattern is the single most useful re.sub feature. Any time the replacement depends on the content of the match (not just rearranging capture groups), reach for it.
What You Learned
- Inline regex pipelines — non-greedy
.+?, ordered passes (bold before italic, images before links). - The state-machine pattern for block-level parsers: state variable, transitions per input line, end-of-document cleanup.
- HTML escaping discipline — escape prose, escape code-block bodies, NEVER escape after generating tags.
re.subwith a callback — when the replacement depends on the match content.- The class-as-state-bag pattern —
self._in_list,self._in_code,self._parainstead of dragging four variables through every helper.
You just built a tiny interpreter. It tokenises (regex), keeps state (the flags and the paragraph buffer), and emits output (the HTML lines). Compilers, template engines, and config parsers all share this shape — once you've written one, every other one feels familiar.
That's the Projects track for now. From here, the next levels build out web (Flask, FastAPI), data (pandas, numpy), and async topics — each project will use the foundations you've practiced across these 13 builds.
Practice this
on practicepython.inShort exercises that run in your browser and tell you what your code actually did, not just whether a test passed.