Web Scraping with requests + BeautifulSoup
1 · The lesson
readScraping is two problems stacked on top of each other: fetch the page (HTTP) and find the data inside it (HTML parsing). Most tutorials nail one half and hand-wave the other. Polite, robust scrapers need both — and they need to behave well enough that the site's admin doesn't notice you exist.
The mental model: you're writing a small, courteous robot. It pretends to be a browser, asks for pages, waits its turn, extracts the bits you want, and saves them in a structure you can re-use. Twenty minutes from now you'll have one.
1. The Two-Library Stack
The de facto stack hasn't changed in a decade:
requests— fetches the page. Synchronous, friendly API, sane defaults.beautifulsoup4— parses the HTML. Forgiving with messy markup, easy traversal.
pip install requests beautifulsoup4
The simplest end-to-end scrape:
import requests from bs4 import BeautifulSoup r = requests.get("https://example.com") r.raise_for_status() # explode loudly on 4xx/5xx — see Section 3 soup = BeautifulSoup(r.text, "html.parser") print(soup.title.text) # "Example Domain"
Three lines of real work. Everything else in this lesson is making that three-line script reliable, polite, and resilient to the real-world ways pages misbehave.
2. Always Send a Real User-Agent
Out of the box, requests sends User-Agent: python-requests/2.31.0 (or whichever). Many sites 403 that header on sight — it screams "bot, possibly an annoying one". The fix is one line:
HEADERS = {
"User-Agent": "Mozilla/5.0 (compatible; HedyBot/1.0; +https://example.com/bot)",
"Accept-Language": "en-GB,en;q=0.9",
}
r = requests.get("https://example.com", headers=HEADERS, timeout=10) setup added so this can run · defines requests
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) requests = _AutoMock('requests')
Two principles at play:
- Identify yourself honestly. Putting a contact URL in the UA is good etiquette — if your scraper misbehaves, the admin can email you instead of nuking your IP.
- Look enough like a browser to not get blocked. A bare
Mozilla/5.0is the bare minimum for sites with naive bot filters.
Some sites need the full browser UA string. You can copy yours from navigator.userAgent in DevTools. Don't pretend to be Googlebot — that's spoofing a known crawler and will get you blacklisted faster than honesty would.
3. Check the Status Code — Always
r = requests.get(url, headers=HEADERS, timeout=10) # Option A — let it raise on any 4xx/5xx r.raise_for_status() # Option B — branch explicitly when you want to handle some codes if r.status_code == 404: return None if r.status_code >= 500: raise RuntimeError("server problem; try again later") r.raise_for_status()
The classic bug: you forget the check, parse the response body, and BeautifulSoup cheerfully parses an HTML error page. Your selectors return nothing, your code thinks the site has no data, and you spend an hour debugging the selector instead of noticing the 503.
r.raise_for_status() raises requests.HTTPError on any non-2xx. Combine with a try/except if you want to log-and-continue rather than crash. See Exceptions for how to do that well.
4. Parsing with BeautifulSoup — Find, Find All, Select
BeautifulSoup gives you two ways to locate elements: the find / find_all API and CSS selectors via select. Both are useful; selectors are usually more concise.
from bs4 import BeautifulSoup html = """ <html><body> <div class="quote" data-id="1"> <span class="text">The unexamined life is not worth living.</span> <span class="author">Socrates</span> <a class="tag" href="/tag/wisdom">wisdom</a> <a class="tag" href="/tag/life">life</a> </div> </body></html> """ soup = BeautifulSoup(html, "html.parser") # find / find_all — by tag name + attrs first = soup.find("span", class_="text") print(first.text) # The unexamined life... quotes = soup.find_all("div", class_="quote") # list of matching tags # CSS selectors — concise, familiar from web work print(soup.select_one("div.quote span.text").text) for a in soup.select("div.quote a.tag"): print(a.text, "→", a.get("href"))
The handful of accessors you'll use over and over:
| Accessor | What it does |
|---|---|
tag.text | All text inside the tag, joined. |
tag.get("href") | Read an attribute; returns None if missing. |
tag["href"] | Same, but raises KeyError if missing. Prefer .get(). |
tag.parent | Walk up one level. |
tag.children | Iterate direct children (mixes tags and NavigableStrings). |
tag.find_next_sibling("li") | Next <li> sibling. |
Use class_= (with the underscore) in find calls — class is a Python keyword. With selectors, the dotted form (.quote) is the normal CSS.
5. Build Robust Selectors — Not Brittle Ones
The biggest reason scrapers break in production: someone changed a CSS class name and your selector silently returned None.
Prefer, in this order:
1. Semantic tags + structure — article > h2, main .entry-title.
2. data-* attributes — these are usually stable hooks the site puts in for its own JS or tests: [data-testid="price"].
3. IDs — usually stable but sometimes generated (e.g. user-7821).
4. Class names — last resort. Especially avoid utility classes (text-sm flex-1) — those change constantly.
# Brittle — relies on a generated class soup price = soup.select_one(".sc-1a3b2c4d-5 .x9k").text # Better — uses a semantic attribute likely added on purpose price = soup.select_one("[data-testid='product-price']").text
setup added so this can run · defines soup
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) soup = _AutoMock('soup')
When the data isn't tagged semantically, your fallback is the structural anchor: find a stable parent, then navigate down. "The <dt> that says 'Price', then the next <dd>":
dt = soup.find("dt", string="Price") price = dt.find_next_sibling("dd").text.strip()
setup added so this can run · defines soup
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) soup = _AutoMock('soup')
This pattern survives a lot of CSS refactors because it reads the page the way a human does.
6. Pagination — Don't Stop at Page One
Two flavours of pagination dominate.
Numbered pages — easy:
def scrape_pages(base_url, max_pages=10): results = [] for page in range(1, max_pages + 1): r = requests.get(f"{base_url}?page={page}", headers=HEADERS, timeout=10) if r.status_code == 404: break # past the last page r.raise_for_status() soup = BeautifulSoup(r.text, "html.parser") items = soup.select("article") if not items: break # empty page — done results.extend(extract(item) for item in items) time.sleep(1) # be polite — see Section 8 return results
setup added so this can run · defines BeautifulSoup, requests, HEADERS, time, extract
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def BeautifulSoup(*_a, **_kw): print('-> BeautifulSoup() called') return _AutoMock('BeautifulSoup()') requests = _AutoMock('requests') HEADERS = _AutoMock('HEADERS') time = _AutoMock('time') def extract(*_a, **_kw): print('-> extract() called') return _AutoMock('extract()')
Next-link pagination — follow the next anchor until it disappears:
url = "https://example.com/blog" while url: r = requests.get(url, headers=HEADERS, timeout=10) r.raise_for_status() soup = BeautifulSoup(r.text, "html.parser") yield from extract_items(soup) next_link = soup.select_one("a[rel='next']") url = next_link["href"] if next_link else None time.sleep(1)
Always cap iterations with max_pages or a sanity counter. The day the site's pagination develops a bug and links page N → page N, your scraper without a cap loops forever.
7. Respect robots.txt
robots.txt is the site's machine-readable "please don't scrape these paths". You're not legally bound by it everywhere, but ignoring it is rude and often a Terms-of-Service violation. The stdlib parses it for you:
from urllib.robotparser import RobotFileParser rp = RobotFileParser() rp.set_url("https://example.com/robots.txt") rp.read() UA = "HedyBot/1.0" if rp.can_fetch(UA, "https://example.com/private/"): fetch(...) else: print("denied by robots.txt")
setup added so this can run · defines fetch
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def fetch(*_a, **_kw): print('-> fetch() called') return _AutoMock('fetch()')
Cache the parsed result per domain — re-fetching robots.txt before every request is wasteful and itself anti-social. Check once per scraper run; refresh hourly for long-running jobs.
8. Rate Limiting and Courtesy
A scraper firing 100 requests per second at a small site looks like a DoS — because it functionally is one. Three lines of time.sleep is the difference between "got the data" and "got banned".
import time for url in urls: fetch(url) time.sleep(0.5) # 2 req/sec, polite default
setup added so this can run · defines urls, fetch
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) urls = ["alpha", "beta", "gamma"] def fetch(*_a, **_kw): print('-> fetch() called') return _AutoMock('fetch()')
Better: rate-limit per domain. If you scrape example.com and other.org in parallel, each gets its own budget:
from collections import defaultdict import time last_request = defaultdict(float) MIN_GAP = 0.5 # seconds between hits per domain def polite_get(url, headers=HEADERS): from urllib.parse import urlparse domain = urlparse(url).netloc wait = MIN_GAP - (time.monotonic() - last_request[domain]) if wait > 0: time.sleep(wait) last_request[domain] = time.monotonic() return requests.get(url, headers=headers, timeout=10)
setup added so this can run · defines HEADERS, requests
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) HEADERS = _AutoMock('HEADERS') requests = _AutoMock('requests')
For larger scrapers, look at the ratelimit package or implement a token bucket. For a one-off, time.sleep is fine.
9. Retries with Exponential Backoff
Networks fail. Servers blink. Your scraper shouldn't die because of one flaky 503.
The hand-rolled version, perfectly adequate for most jobs:
import time import requests def fetch_with_retries(url, max_attempts=4): for attempt in range(max_attempts): try: r = requests.get(url, headers=HEADERS, timeout=10) if r.status_code < 500: return r # 2xx / 3xx / 4xx — don't retry client errors except requests.RequestException: pass sleep = 2 ** attempt # 1s, 2s, 4s, 8s time.sleep(sleep) raise RuntimeError(f"gave up after {max_attempts} attempts: {url}")
setup added so this can run · defines HEADERS
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) HEADERS = _AutoMock('HEADERS')
The production-grade version uses urllib3.util.Retry mounted on a session:
from requests.adapters import HTTPAdapter from urllib3.util.retry import Retry retry = Retry( total=4, backoff_factor=1, # 0s, 2s, 4s, 8s between retries status_forcelist=[429, 500, 502, 503, 504], allowed_methods=["GET", "HEAD"], ) session = requests.Session() session.mount("https://", HTTPAdapter(max_retries=retry))
setup added so this can run · defines requests
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) requests = _AutoMock('requests')
Retry 5xx and 429 (rate-limited). Don't retry 4xx — those are your fault and won't fix themselves. Cap total attempts so a permanent outage doesn't loop forever.
10. Use a Session — Cookies, Pooling, Defaults
For more than two requests to the same host, use a requests.Session. It reuses the underlying TCP connection (huge speedup), persists cookies (necessary for anything login-gated), and lets you set headers in one place.
with requests.Session() as s: s.headers.update(HEADERS) # applied to every request s.get("https://example.com/login") # cookies stored automatically r = s.get("https://example.com/dashboard") r.raise_for_status() process(r.text)
setup added so this can run · defines HEADERS, process, requests
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) HEADERS = _AutoMock('HEADERS') def process(*_a, **_kw): print('-> process() called') return _AutoMock('process()') requests = _AutoMock('requests')
The connection-pool win alone is worth it on big scrapes — a fresh TCP handshake per request adds tens of milliseconds and CPU. Combine with the Retry adapter from Section 9 and you have a solid foundation.
11. When the Page Is JavaScript-Rendered
requests fetches the initial HTML response. If the page loads its real content via JavaScript after that (think most React/Vue/Svelte SPAs), your soup will be near-empty — just a loading shell.
Two signs you've hit this wall:
view-source:in your browser shows mostly empty<div id="root">containers.- The data you see in DevTools doesn't appear in
requests.get(url).text.
The escape hatches:
- Find the underlying API. Open DevTools → Network → XHR. Often the page is fetching its data from a JSON endpoint you can hit directly — much cleaner than scraping the rendered DOM. (See APIs.)
- Use a real browser via Playwright or Selenium.
playwrightis the modern choice —await page.goto(url), wait for the selector, then read the rendered HTML. Slower, heavier, but reliable.
requests + BeautifulSoup covers maybe 70% of public scraping. The other 30% needs a headless browser, and pretending otherwise leads to brittle, unreliable scripts.
12. Saving the Output
Once you've extracted data, dump it somewhere structured. JSONL is the streaming-friendly default; CSV is for spreadsheet-bound data; HTML archive is for "I might need to re-parse later":
import json from pathlib import Path out = Path("quotes.jsonl") with out.open("w", encoding="utf-8") as f: for record in records: f.write(json.dumps(record, ensure_ascii=False) + "\n")
setup added so this can run · defines records
records = ["alpha", "beta", "gamma"]
For tabular output and the full set of pitfalls (encoding, newline modes, streaming), see CSV & JSON. For archiving raw HTML, path.write_bytes(r.content) and you'll have something you can re-parse offline forever.
13. Scrape Responsibly — Legal and Ethical
Scraping sits in legal grey areas that differ by jurisdiction. A short, practical checklist:
- Read the Terms of Service. Many explicitly forbid automated access. Violating them rarely lands you in court, but it can get your account or IP banned.
- Respect robots.txt (Section 7).
- Don't redistribute copyrighted content wholesale — extracting facts is generally OK; reposting articles is not.
- Avoid personal data. Names, emails, addresses scraped from public pages are still personal data under GDPR and similar laws.
- Rate-limit yourself so the site barely notices you (Section 8). Costing a small site real money in bandwidth is the fastest way to get blocked.
- Cache aggressively. If you re-parse the same page in development, save the HTML once and read it from disk on subsequent runs. Your iteration loop will be faster and you'll be a better citizen.
When in doubt: would you be embarrassed if the site's admin saw your traffic in their logs? If yes, slow it down or stop.
14. Common Mistakes
1. No User-Agent set
Default python-requests/2.x UA → 403 on a lot of real-world sites. One line of headers prevents it.
2. No rate limiting
Firing requests as fast as possible gets your IP banned, hurts the site, and tells the world you're an amateur. time.sleep(0.5) minimum.
3. Brittle selectors based on utility classes.sc-1a3b2c4d-5 will change next deployment. Anchor on semantics, data-*, or stable IDs first.
4. Not checking r.status_code
Parsing an HTML error page silently produces "no results" bugs that take hours to diagnose. r.raise_for_status() or branch explicitly.
5. Hardcoding cookies/session state
Pasting a session cookie into your script works for ten minutes, then expires. Use requests.Session and log in properly, or use the site's API.
6. except: pass around the whole loop
You'll catch the KeyboardInterrupt you sent to stop the runaway scraper. Catch specific exceptions — requests.RequestException for network, AttributeError for missing selectors — and log them. See Exceptions.
7. No timeoutrequests.get(url) with no timeout= can hang forever if the server stops responding mid-stream. Always set one.
8. Re-fetching during development
You're tweaking your selector and hammering the site 50 times with each save. Cache the HTML to disk once; iterate locally. Faster and polite.
🎯 Your Turn — Scrape quotes.toscrape.com
https://quotes.toscrape.com/ is a sandbox site built exactly for this. It has numbered pages at /page/1/, /page/2/, …, and structured quote markup. Write scrape_quotes(url, max_pages=5) that returns a list of dicts:
[
{"quote": "The world as we have created it...",
"author": "Albert Einstein",
"tags": ["change", "deep-thoughts", "thinking", "world"]},
...
]Constraints:
- One-second sleep between page fetches.
- Send a real User-Agent.
- Stop early when a page has zero quotes (you've gone past the last one).
- Set a timeout on every request.
Skeleton:
import time import requests from bs4 import BeautifulSoup HEADERS = {"User-Agent": "Mozilla/5.0 (compatible; LearnerBot/1.0)"} def scrape_quotes(base_url="https://quotes.toscrape.com", max_pages=5): results = [] for page in range(1, max_pages + 1): url = f"{base_url}/page/{page}/" # TODO 1: GET with headers and timeout; check status # TODO 2: parse with BeautifulSoup # TODO 3: find each quote block; extract text, author, tags # TODO 4: break early if the page yields no quotes # TODO 5: sleep 1 second between pages ... return results
Hint 1 — Inspect the markup
Each quote is wrapped in<div class="quote">. Inside: <span class="text"> for the quote, <small class="author"> for the author, and <a class="tag"> elements for each tag. Use select with CSS selectors — concise and readable.
Hint 2 — Stripping the curly quotes
The site wraps quotes in U+201C / U+201D curly quotes..text.strip("“”") removes them. Or leave them in — it's data, you can clean later.
Show full solution
import time import requests from bs4 import BeautifulSoup HEADERS = {"User-Agent": "Mozilla/5.0 (compatible; LearnerBot/1.0)"} def scrape_quotes(base_url="https://quotes.toscrape.com", max_pages=5): """Scrape quotes from quotes.toscrape.com, politely.""" results = [] for page in range(1, max_pages + 1): url = f"{base_url}/page/{page}/" r = requests.get(url, headers=HEADERS, timeout=10) r.raise_for_status() soup = BeautifulSoup(r.text, "html.parser") blocks = soup.select("div.quote") if not blocks: break # past the last page for block in blocks: results.append({ "quote": block.select_one("span.text").text.strip("“”"), "author": block.select_one("small.author").text, "tags": [a.text for a in block.select("a.tag")], }) time.sleep(1) # be polite return results # Demo quotes = scrape_quotes(max_pages=2) print(f"scraped {len(quotes)} quotes") for q in quotes[:2]: print(f"- {q['author']}: {q['quote'][:60]}... [{', '.join(q['tags'])}]")
What you did:
- Sent a real User-Agent in every request — no anonymous
python-requestssignature. - Checked the response with
raise_for_status()so a 5xx error fails loudly instead of producing empty results. - Set a
timeout=10so a stalled server doesn't hang your script forever. - Used CSS selectors with
selectandselect_one— concise and readable. - Broke early when a page has zero quotes — the natural end-of-data signal.
- Slept one second between pages — the site barely notices you exist.
The structure is the same shape every static scraper takes: loop, fetch, check, parse, extract, sleep. Once you internalise that rhythm, scraping any static site is mostly figuring out the right selectors.
For a real production scraper you'd also: wrap retries with urllib3.util.Retry, store results incrementally (so a crash doesn't lose everything), and cache fetched HTML to disk during development. All shown in the lesson above — you've got everything you need.
What You Learned
- The stack:
requestsfor fetching,BeautifulSoupfor parsing. Three lines do the simple case. - Always set a User-Agent and a
timeout. Always checkr.status_codeor callraise_for_status(). - Selectors:
.find(),.find_all(),.select(),.select_one(). Navigate with.text,.get("href"),.parent. - Robust selectors prefer semantic tags,
data-*attributes, and structural anchors over utility class names. - Pagination: loop
?page=Nuntil empty, or followa[rel='next']until missing. Cap iterations. - Politeness: respect
robots.txt, sleep between requests, rate-limit per domain. - Retries:
urllib3.util.Retrymounted on aSession, or hand-rolled exponential backoff. Retry 5xx and 429, never 4xx. - Sessions reuse connections and persist cookies — use them for anything multi-request.
- JS-rendered pages need Playwright/Selenium, or look for the underlying API (APIs).
- Save structured output as JSON/CSV/JSONL — see CSV & JSON.
- Scrape responsibly — ToS, copyright, personal data, rate limits. Be the bot you'd want to host.
Next: REST APIs — when there's a JSON endpoint, scraping is the wrong tool. Hit the API instead.