PythonMastery
intermediate 20 min read · lesson 1 of 12 in Python How-To

Web Scraping with requests + BeautifulSoup

1 · The lesson

read

Scraping is two problems stacked on top of each other: fetch the page (HTTP) and find the data inside it (HTML parsing). Most tutorials nail one half and hand-wave the other. Polite, robust scrapers need both — and they need to behave well enough that the site's admin doesn't notice you exist.

The mental model: you're writing a small, courteous robot. It pretends to be a browser, asks for pages, waits its turn, extracts the bits you want, and saves them in a structure you can re-use. Twenty minutes from now you'll have one.


1. The Two-Library Stack

The de facto stack hasn't changed in a decade:

  • requests — fetches the page. Synchronous, friendly API, sane defaults.
  • beautifulsoup4 — parses the HTML. Forgiving with messy markup, easy traversal.
bash
pip install requests beautifulsoup4

The simplest end-to-end scrape:

python
import requests
from bs4 import BeautifulSoup

r = requests.get("https://example.com")
r.raise_for_status()                 # explode loudly on 4xx/5xx — see Section 3
soup = BeautifulSoup(r.text, "html.parser")
print(soup.title.text)               # "Example Domain"

Three lines of real work. Everything else in this lesson is making that three-line script reliable, polite, and resilient to the real-world ways pages misbehave.


2. Always Send a Real User-Agent

Out of the box, requests sends User-Agent: python-requests/2.31.0 (or whichever). Many sites 403 that header on sight — it screams "bot, possibly an annoying one". The fix is one line:

python
HEADERS = {
    "User-Agent": "Mozilla/5.0 (compatible; HedyBot/1.0; +https://example.com/bot)",
    "Accept-Language": "en-GB,en;q=0.9",
}

r = requests.get("https://example.com", headers=HEADERS, timeout=10)
+ setup added so this can run · defines requests
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

requests = _AutoMock('requests')

Two principles at play:

  • Identify yourself honestly. Putting a contact URL in the UA is good etiquette — if your scraper misbehaves, the admin can email you instead of nuking your IP.
  • Look enough like a browser to not get blocked. A bare Mozilla/5.0 is the bare minimum for sites with naive bot filters.

Some sites need the full browser UA string. You can copy yours from navigator.userAgent in DevTools. Don't pretend to be Googlebot — that's spoofing a known crawler and will get you blacklisted faster than honesty would.


3. Check the Status Code — Always

python
r = requests.get(url, headers=HEADERS, timeout=10)

# Option A — let it raise on any 4xx/5xx
r.raise_for_status()

# Option B — branch explicitly when you want to handle some codes
if r.status_code == 404:
    return None
if r.status_code >= 500:
    raise RuntimeError("server problem; try again later")
r.raise_for_status()

The classic bug: you forget the check, parse the response body, and BeautifulSoup cheerfully parses an HTML error page. Your selectors return nothing, your code thinks the site has no data, and you spend an hour debugging the selector instead of noticing the 503.

r.raise_for_status() raises requests.HTTPError on any non-2xx. Combine with a try/except if you want to log-and-continue rather than crash. See Exceptions for how to do that well.


4. Parsing with BeautifulSoup — Find, Find All, Select

BeautifulSoup gives you two ways to locate elements: the find / find_all API and CSS selectors via select. Both are useful; selectors are usually more concise.

python
from bs4 import BeautifulSoup

html = """
<html><body>
  <div class="quote" data-id="1">
    <span class="text">The unexamined life is not worth living.</span>
    <span class="author">Socrates</span>
    <a class="tag" href="/tag/wisdom">wisdom</a>
    <a class="tag" href="/tag/life">life</a>
  </div>
</body></html>
"""

soup = BeautifulSoup(html, "html.parser")

# find / find_all — by tag name + attrs
first = soup.find("span", class_="text")
print(first.text)                                    # The unexamined life...
quotes = soup.find_all("div", class_="quote")        # list of matching tags

# CSS selectors — concise, familiar from web work
print(soup.select_one("div.quote span.text").text)
for a in soup.select("div.quote a.tag"):
    print(a.text, "→", a.get("href"))

The handful of accessors you'll use over and over:

AccessorWhat it does
tag.textAll text inside the tag, joined.
tag.get("href")Read an attribute; returns None if missing.
tag["href"]Same, but raises KeyError if missing. Prefer .get().
tag.parentWalk up one level.
tag.childrenIterate direct children (mixes tags and NavigableStrings).
tag.find_next_sibling("li")Next <li> sibling.

Use class_= (with the underscore) in find calls — class is a Python keyword. With selectors, the dotted form (.quote) is the normal CSS.


5. Build Robust Selectors — Not Brittle Ones

The biggest reason scrapers break in production: someone changed a CSS class name and your selector silently returned None.

Prefer, in this order:

1. Semantic tags + structure — article > h2, main .entry-title.
2. data-* attributes — these are usually stable hooks the site puts in for its own JS or tests: [data-testid="price"].
3. IDs — usually stable but sometimes generated (e.g. user-7821).
4. Class names — last resort. Especially avoid utility classes (text-sm flex-1) — those change constantly.

python
# Brittle — relies on a generated class soup
price = soup.select_one(".sc-1a3b2c4d-5 .x9k").text

# Better — uses a semantic attribute likely added on purpose
price = soup.select_one("[data-testid='product-price']").text
+ setup added so this can run · defines soup
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

soup = _AutoMock('soup')

When the data isn't tagged semantically, your fallback is the structural anchor: find a stable parent, then navigate down. "The <dt> that says 'Price', then the next <dd>":

python
dt = soup.find("dt", string="Price")
price = dt.find_next_sibling("dd").text.strip()
+ setup added so this can run · defines soup
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

soup = _AutoMock('soup')

This pattern survives a lot of CSS refactors because it reads the page the way a human does.


6. Pagination — Don't Stop at Page One

Two flavours of pagination dominate.

Numbered pages — easy:

python
def scrape_pages(base_url, max_pages=10):
    results = []
    for page in range(1, max_pages + 1):
        r = requests.get(f"{base_url}?page={page}", headers=HEADERS, timeout=10)
        if r.status_code == 404:
            break                                # past the last page
        r.raise_for_status()
        soup = BeautifulSoup(r.text, "html.parser")
        items = soup.select("article")
        if not items:
            break                                # empty page — done
        results.extend(extract(item) for item in items)
        time.sleep(1)                            # be polite — see Section 8
    return results
+ setup added so this can run · defines BeautifulSoup, requests, HEADERS, time, extract
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def BeautifulSoup(*_a, **_kw):
    print('-> BeautifulSoup() called')
    return _AutoMock('BeautifulSoup()')
requests = _AutoMock('requests')
HEADERS = _AutoMock('HEADERS')
time = _AutoMock('time')
def extract(*_a, **_kw):
    print('-> extract() called')
    return _AutoMock('extract()')

Next-link pagination — follow the next anchor until it disappears:

python
url = "https://example.com/blog"
while url:
    r = requests.get(url, headers=HEADERS, timeout=10)
    r.raise_for_status()
    soup = BeautifulSoup(r.text, "html.parser")
    yield from extract_items(soup)
    next_link = soup.select_one("a[rel='next']")
    url = next_link["href"] if next_link else None
    time.sleep(1)

Always cap iterations with max_pages or a sanity counter. The day the site's pagination develops a bug and links page N → page N, your scraper without a cap loops forever.


7. Respect robots.txt

robots.txt is the site's machine-readable "please don't scrape these paths". You're not legally bound by it everywhere, but ignoring it is rude and often a Terms-of-Service violation. The stdlib parses it for you:

python
from urllib.robotparser import RobotFileParser

rp = RobotFileParser()
rp.set_url("https://example.com/robots.txt")
rp.read()

UA = "HedyBot/1.0"
if rp.can_fetch(UA, "https://example.com/private/"):
    fetch(...)
else:
    print("denied by robots.txt")
+ setup added so this can run · defines fetch
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def fetch(*_a, **_kw):
    print('-> fetch() called')
    return _AutoMock('fetch()')

Cache the parsed result per domain — re-fetching robots.txt before every request is wasteful and itself anti-social. Check once per scraper run; refresh hourly for long-running jobs.


8. Rate Limiting and Courtesy

A scraper firing 100 requests per second at a small site looks like a DoS — because it functionally is one. Three lines of time.sleep is the difference between "got the data" and "got banned".

python
import time

for url in urls:
    fetch(url)
    time.sleep(0.5)                              # 2 req/sec, polite default
+ setup added so this can run · defines urls, fetch
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

urls = ["alpha", "beta", "gamma"]
def fetch(*_a, **_kw):
    print('-> fetch() called')
    return _AutoMock('fetch()')

Better: rate-limit per domain. If you scrape example.com and other.org in parallel, each gets its own budget:

python
from collections import defaultdict
import time

last_request = defaultdict(float)
MIN_GAP = 0.5                                    # seconds between hits per domain

def polite_get(url, headers=HEADERS):
    from urllib.parse import urlparse
    domain = urlparse(url).netloc
    wait = MIN_GAP - (time.monotonic() - last_request[domain])
    if wait > 0:
        time.sleep(wait)
    last_request[domain] = time.monotonic()
    return requests.get(url, headers=headers, timeout=10)
+ setup added so this can run · defines HEADERS, requests
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

HEADERS = _AutoMock('HEADERS')
requests = _AutoMock('requests')

For larger scrapers, look at the ratelimit package or implement a token bucket. For a one-off, time.sleep is fine.


9. Retries with Exponential Backoff

Networks fail. Servers blink. Your scraper shouldn't die because of one flaky 503.

The hand-rolled version, perfectly adequate for most jobs:

python
import time
import requests

def fetch_with_retries(url, max_attempts=4):
    for attempt in range(max_attempts):
        try:
            r = requests.get(url, headers=HEADERS, timeout=10)
            if r.status_code < 500:
                return r                         # 2xx / 3xx / 4xx — don't retry client errors
        except requests.RequestException:
            pass
        sleep = 2 ** attempt                     # 1s, 2s, 4s, 8s
        time.sleep(sleep)
    raise RuntimeError(f"gave up after {max_attempts} attempts: {url}")
+ setup added so this can run · defines HEADERS
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

HEADERS = _AutoMock('HEADERS')

The production-grade version uses urllib3.util.Retry mounted on a session:

python
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

retry = Retry(
    total=4,
    backoff_factor=1,                            # 0s, 2s, 4s, 8s between retries
    status_forcelist=[429, 500, 502, 503, 504],
    allowed_methods=["GET", "HEAD"],
)
session = requests.Session()
session.mount("https://", HTTPAdapter(max_retries=retry))
+ setup added so this can run · defines requests
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

requests = _AutoMock('requests')

Retry 5xx and 429 (rate-limited). Don't retry 4xx — those are your fault and won't fix themselves. Cap total attempts so a permanent outage doesn't loop forever.


10. Use a Session — Cookies, Pooling, Defaults

For more than two requests to the same host, use a requests.Session. It reuses the underlying TCP connection (huge speedup), persists cookies (necessary for anything login-gated), and lets you set headers in one place.

python
with requests.Session() as s:
    s.headers.update(HEADERS)                    # applied to every request
    s.get("https://example.com/login")           # cookies stored automatically
    r = s.get("https://example.com/dashboard")
    r.raise_for_status()
    process(r.text)
+ setup added so this can run · defines HEADERS, process, requests
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

HEADERS = _AutoMock('HEADERS')
def process(*_a, **_kw):
    print('-> process() called')
    return _AutoMock('process()')
requests = _AutoMock('requests')

The connection-pool win alone is worth it on big scrapes — a fresh TCP handshake per request adds tens of milliseconds and CPU. Combine with the Retry adapter from Section 9 and you have a solid foundation.


11. When the Page Is JavaScript-Rendered

requests fetches the initial HTML response. If the page loads its real content via JavaScript after that (think most React/Vue/Svelte SPAs), your soup will be near-empty — just a loading shell.

Two signs you've hit this wall:

  • view-source: in your browser shows mostly empty <div id="root"> containers.
  • The data you see in DevTools doesn't appear in requests.get(url).text.

The escape hatches:

  • Find the underlying API. Open DevTools → Network → XHR. Often the page is fetching its data from a JSON endpoint you can hit directly — much cleaner than scraping the rendered DOM. (See APIs.)
  • Use a real browser via Playwright or Selenium. playwright is the modern choice — await page.goto(url), wait for the selector, then read the rendered HTML. Slower, heavier, but reliable.

requests + BeautifulSoup covers maybe 70% of public scraping. The other 30% needs a headless browser, and pretending otherwise leads to brittle, unreliable scripts.


12. Saving the Output

Once you've extracted data, dump it somewhere structured. JSONL is the streaming-friendly default; CSV is for spreadsheet-bound data; HTML archive is for "I might need to re-parse later":

python
import json
from pathlib import Path

out = Path("quotes.jsonl")
with out.open("w", encoding="utf-8") as f:
    for record in records:
        f.write(json.dumps(record, ensure_ascii=False) + "\n")
+ setup added so this can run · defines records
records = ["alpha", "beta", "gamma"]

For tabular output and the full set of pitfalls (encoding, newline modes, streaming), see CSV & JSON. For archiving raw HTML, path.write_bytes(r.content) and you'll have something you can re-parse offline forever.


Scraping sits in legal grey areas that differ by jurisdiction. A short, practical checklist:

  • Read the Terms of Service. Many explicitly forbid automated access. Violating them rarely lands you in court, but it can get your account or IP banned.
  • Respect robots.txt (Section 7).
  • Don't redistribute copyrighted content wholesale — extracting facts is generally OK; reposting articles is not.
  • Avoid personal data. Names, emails, addresses scraped from public pages are still personal data under GDPR and similar laws.
  • Rate-limit yourself so the site barely notices you (Section 8). Costing a small site real money in bandwidth is the fastest way to get blocked.
  • Cache aggressively. If you re-parse the same page in development, save the HTML once and read it from disk on subsequent runs. Your iteration loop will be faster and you'll be a better citizen.

When in doubt: would you be embarrassed if the site's admin saw your traffic in their logs? If yes, slow it down or stop.


14. Common Mistakes

1. No User-Agent set
Default python-requests/2.x UA → 403 on a lot of real-world sites. One line of headers prevents it.

2. No rate limiting
Firing requests as fast as possible gets your IP banned, hurts the site, and tells the world you're an amateur. time.sleep(0.5) minimum.

3. Brittle selectors based on utility classes
.sc-1a3b2c4d-5 will change next deployment. Anchor on semantics, data-*, or stable IDs first.

4. Not checking r.status_code
Parsing an HTML error page silently produces "no results" bugs that take hours to diagnose. r.raise_for_status() or branch explicitly.

5. Hardcoding cookies/session state
Pasting a session cookie into your script works for ten minutes, then expires. Use requests.Session and log in properly, or use the site's API.

6. except: pass around the whole loop
You'll catch the KeyboardInterrupt you sent to stop the runaway scraper. Catch specific exceptions — requests.RequestException for network, AttributeError for missing selectors — and log them. See Exceptions.

7. No timeout
requests.get(url) with no timeout= can hang forever if the server stops responding mid-stream. Always set one.

8. Re-fetching during development
You're tweaking your selector and hammering the site 50 times with each save. Cache the HTML to disk once; iterate locally. Faster and polite.


🎯 Your Turn — Scrape quotes.toscrape.com

https://quotes.toscrape.com/ is a sandbox site built exactly for this. It has numbered pages at /page/1/, /page/2/, …, and structured quote markup. Write scrape_quotes(url, max_pages=5) that returns a list of dicts:

python
[
    {"quote": "The world as we have created it...",
     "author": "Albert Einstein",
     "tags": ["change", "deep-thoughts", "thinking", "world"]},
    ...
]

Constraints:

  • One-second sleep between page fetches.
  • Send a real User-Agent.
  • Stop early when a page has zero quotes (you've gone past the last one).
  • Set a timeout on every request.

Skeleton:

python
import time
import requests
from bs4 import BeautifulSoup

HEADERS = {"User-Agent": "Mozilla/5.0 (compatible; LearnerBot/1.0)"}

def scrape_quotes(base_url="https://quotes.toscrape.com", max_pages=5):
    results = []
    for page in range(1, max_pages + 1):
        url = f"{base_url}/page/{page}/"
        # TODO 1: GET with headers and timeout; check status
        # TODO 2: parse with BeautifulSoup
        # TODO 3: find each quote block; extract text, author, tags
        # TODO 4: break early if the page yields no quotes
        # TODO 5: sleep 1 second between pages
        ...
    return results
Hint 1 — Inspect the markup Each quote is wrapped in <div class="quote">. Inside: <span class="text"> for the quote, <small class="author"> for the author, and <a class="tag"> elements for each tag. Use select with CSS selectors — concise and readable.
Hint 2 — Stripping the curly quotes The site wraps quotes in U+201C / U+201D curly quotes. .text.strip("“”") removes them. Or leave them in — it's data, you can clean later.
Show full solution
python
import time
import requests
from bs4 import BeautifulSoup

HEADERS = {"User-Agent": "Mozilla/5.0 (compatible; LearnerBot/1.0)"}

def scrape_quotes(base_url="https://quotes.toscrape.com", max_pages=5):
    """Scrape quotes from quotes.toscrape.com, politely."""
    results = []
    for page in range(1, max_pages + 1):
        url = f"{base_url}/page/{page}/"
        r = requests.get(url, headers=HEADERS, timeout=10)
        r.raise_for_status()
        soup = BeautifulSoup(r.text, "html.parser")
        blocks = soup.select("div.quote")
        if not blocks:
            break                                       # past the last page
        for block in blocks:
            results.append({
                "quote":  block.select_one("span.text").text.strip("“”"),
                "author": block.select_one("small.author").text,
                "tags":   [a.text for a in block.select("a.tag")],
            })
        time.sleep(1)                                   # be polite
    return results


# Demo
quotes = scrape_quotes(max_pages=2)
print(f"scraped {len(quotes)} quotes")
for q in quotes[:2]:
    print(f"- {q['author']}: {q['quote'][:60]}... [{', '.join(q['tags'])}]")

What you did:

  • Sent a real User-Agent in every request — no anonymous python-requests signature.
  • Checked the response with raise_for_status() so a 5xx error fails loudly instead of producing empty results.
  • Set a timeout=10 so a stalled server doesn't hang your script forever.
  • Used CSS selectors with select and select_one — concise and readable.
  • Broke early when a page has zero quotes — the natural end-of-data signal.
  • Slept one second between pages — the site barely notices you exist.

The structure is the same shape every static scraper takes: loop, fetch, check, parse, extract, sleep. Once you internalise that rhythm, scraping any static site is mostly figuring out the right selectors.

For a real production scraper you'd also: wrap retries with urllib3.util.Retry, store results incrementally (so a crash doesn't lose everything), and cache fetched HTML to disk during development. All shown in the lesson above — you've got everything you need.


What You Learned

  • The stack: requests for fetching, BeautifulSoup for parsing. Three lines do the simple case.
  • Always set a User-Agent and a timeout. Always check r.status_code or call raise_for_status().
  • Selectors: .find(), .find_all(), .select(), .select_one(). Navigate with .text, .get("href"), .parent.
  • Robust selectors prefer semantic tags, data-* attributes, and structural anchors over utility class names.
  • Pagination: loop ?page=N until empty, or follow a[rel='next'] until missing. Cap iterations.
  • Politeness: respect robots.txt, sleep between requests, rate-limit per domain.
  • Retries: urllib3.util.Retry mounted on a Session, or hand-rolled exponential backoff. Retry 5xx and 429, never 4xx.
  • Sessions reuse connections and persist cookies — use them for anything multi-request.
  • JS-rendered pages need Playwright/Selenium, or look for the underlying API (APIs).
  • Save structured output as JSON/CSV/JSONL — see CSV & JSON.
  • Scrape responsibly — ToS, copyright, personal data, rate limits. Be the bot you'd want to host.

Next: REST APIs — when there's a JSON endpoint, scraping is the wrong tool. Hit the API instead.