pip & the Standard Library: Batteries Included
1 · The lesson
readPython has two layers of code beyond what you write yourself:
1. The standard library — modules that ship with Python. No install required.
2. PyPI (Python Package Index) — 500,000+ community packages you install with pip.
This lesson is a tour of the most useful bits of both. You won't memorize anything here; the point is to know what exists, so when you have a problem you can think "I bet there's a module for that" and Google the right name.
1. The Standard Library — "Batteries Included"
When you import something without installing it first, you're using the standard library. It comes with Python.
Here's the working developer's shortlist — the modules you'll touch most often in real Python code.
Files & paths
import os # operating system interactions import sys # interpreter / script-level info from pathlib import Path # modern path manipulation import shutil # high-level file operations (copy, move, delete trees) import glob # find files by pattern import tempfile # temporary files and directories
Time & dates
import datetime # dates, times, durations import time # low-level time, sleep, timestamps import calendar # calendar math import zoneinfo # timezone handling (Python 3.9+)
Data & formats
import json # JSON read/write import csv # CSV read/write import sqlite3 # built-in SQL database, zero install import pickle # Python's own serialization (use with caution) import base64 # encoding import hashlib # SHA-256, MD5, etc. import uuid # unique IDs
Numbers & math
import math # sqrt, log, trig (covered in Numbers in Depth) import random # random numbers, choices, shuffles import statistics # mean, median, stdev — built in! import decimal # exact decimal arithmetic import fractions # rational numbers (1/3 stays as 1/3)
Text
import re # regex import string # string constants, formatting helpers import textwrap # wrap and indent text import unicodedata # character lookups
Networking & web
import urllib.request # download a URL import http.server # a one-line dev server import socket # low-level networking import smtplib # send email import email # parse/build emails
Collections & data structures
from collections import ( defaultdict, Counter, deque, namedtuple, OrderedDict, ) import itertools # combinatorics, infinite iterators import functools # higher-order helpers (lru_cache, partial) import heapq # priority queue import bisect # binary search in sorted lists
Tools & utilities
import argparse # command-line argument parsing import logging # structured logging import unittest # test framework (or use pytest from PyPI) import subprocess # run other programs import threading # threads import multiprocessing # processes (true parallelism) import asyncio # async/await
Don't try to memorize this — bookmark it. The point is: for almost any beginner problem, the standard library already has it.
2. A Few Examples — "There's a Module for That"
Read a JSON file:
import json data_str = '{"name": "Alice", "age": 30, "skills": ["python", "sql"]}' data = json.loads(data_str) print(data["skills"]) # ['python', 'sql'] # Reverse direction — Python → JSON string out = json.dumps({"x": 1, "y": 2}, indent=2) print(out)
Get the current date and format it:
from datetime import datetime, timedelta now = datetime.now() print(now.strftime("%Y-%m-%d %H:%M")) # e.g. 2026-05-14 09:30 # Tomorrow tomorrow = now + timedelta(days=1) print(tomorrow.strftime("%A")) # the weekday name
Count occurrences of items in a list:
from collections import Counter votes = ["red", "blue", "red", "green", "blue", "red"] tally = Counter(votes) print(tally) # Counter({'red': 3, 'blue': 2, 'green': 1}) print(tally.most_common(2)) # [('red', 3), ('blue', 2)]
Compute a SHA-256 hash:
import hashlib digest = hashlib.sha256(b"hello world").hexdigest() print(digest) # b94d27b9934d3e08a52e52d7da7dabfac484efe37a5380ee9088f7ace2efcde9
Generate a unique ID:
import uuid print(uuid.uuid4()) # e.g. 7c9e6679-7425-40de-944b-e07fc1f90ae7
Five entirely different problems, five built-in modules, zero installs.
3. pip — Installing Third-Party Packages
When the standard library doesn't have what you need, you install from PyPI using pip. (You can't run pip in the browser sandbox, but the commands look like this.)
# Install a single package pip install requests # Install a specific version pip install requests==2.31.0 # Install several at once pip install requests pandas numpy # Upgrade pip install --upgrade requests # List installed packages pip list # Show details about a package pip show requests # Remove pip uninstall requests
The package then becomes importable in your code:
# After `pip install requests` import requests response = requests.get("https://api.github.com") print(response.status_code) # 200
4. The "Top 10" Packages You'll Reach for from PyPI
Once you graduate from pure-standard-library, these are the names you'll meet again and again:
| Package | What it's for |
|---|---|
| requests | HTTP — make API calls, download files. The friendlier urllib.request. |
| numpy | Fast numeric arrays. Foundation of the data stack. |
| pandas | Tabular data. Spreadsheets, CSVs, SQL results — all become DataFrames. |
| matplotlib | Charts and plots. The classic. |
| scikit-learn | Classical machine learning algorithms. |
| flask / fastapi | Build web APIs. |
| beautifulsoup4 | Parse HTML — for web scraping. |
| pillow | Image processing. |
| pytest | The de facto test framework. |
| rich | Beautiful terminal output (colors, tables, progress bars). |
You'll meet each of these in a dedicated lesson later in the catalog.
5. requirements.txt — Saving Your Project's Dependencies
Real projects track their dependencies in a file so anyone can recreate the environment:
# In your project directory: pip freeze > requirements.txt # Someone else (or future-you on a different machine): pip install -r requirements.txt
A requirements.txt looks like:
requests==2.31.0 pandas==2.0.3 numpy==1.25.0
Pinning versions (with ==) is the safest practice — your code keeps working even if upstream packages change.
6. Virtual Environments — The 30-Second Version
If you install packages globally, your projects start fighting each other for versions. The fix is a virtual environment — an isolated Python install per project.
# Create a virtual environment python -m venv .venv # Activate it # macOS/Linux: source .venv/bin/activate # Windows: .venv\Scripts\activate # Install into the venv (now isolated from other projects) pip install requests # Exit when done deactivate
The full treatment lives in the Virtual Environments lesson (intermediate). For now, just know: always use a venv for real projects.
7. The Discovery Loop — How to Find Modules
When you have a problem:
1. Search "python how to X" — first hit is usually a module
2. help(module) in a REPL for a quick overview
3. dir(module) to list everything in a module
4. docs.python.org for the canonical reference
5. PyPI search at pypi.org for third-party packages
Don't write something from scratch until you've checked. Especially for things like dates, CSV parsing, HTTP, or hashing — someone has solved it carefully already.
8. Mistakes You'll Hit
1. Installing packages globally
Eventually two projects need different versions of the same package. Use venv.
2. Reinventing what the stdlib has
Don't write your own CSV parser. Don't write your own datetime arithmetic. Don't write your own hash function. The stdlib versions handle edge cases you don't know about.
3. Pinning versions too loosely or not at all
Without == in requirements.txt, your code can break six months later because a dependency had a breaking change.
4. pip install failing? Try python -m pip install
On some systems, pip points to the wrong Python. python -m pip always uses the matching Python.
Mini-Reference — One-Liners
# Pretty-print a Python value import pprint; pprint.pprint({"a": list(range(20))}) # Time a chunk of code import timeit print(timeit.timeit("'-'.join(str(n) for n in range(100))", number=10000)) # Read a URL (stdlib version, no requests needed) from urllib.request import urlopen # with urlopen("https://example.com") as r: # html = r.read().decode() # Run a shell command and capture output import subprocess # result = subprocess.run(["ls", "-l"], capture_output=True, text=True) # print(result.stdout) # Get an environment variable import os home = os.environ.get("HOME") or os.environ.get("USERPROFILE") print(home)
setup added so this can run · defines
import os # noqa: F401 os.environ.setdefault("HOME", "example-home") os.environ.setdefault("USERPROFILE", "example-userprofile")
🎯 Your Turn — Reach for the Standard Library First
The point of "batteries included" is that a surprising amount of work is already
done for you. This exercise is deliberately easy to write badly and short to
write well.
Given a list of survey responses, produce a small report: the three most common
answers with their counts, and the median response length. Use the standard
library — no pip install, no hand-rolled counting or sorting.
responses = ["yes", "no", "yes", "maybe", "yes", "no", "absolutely"] Top 3 : yes (3), no (2), maybe (1) Median len : 3
Skeleton:
from collections import Counter import statistics def report(responses): # TODO 1: use Counter to get the three most common answers # TODO 2: use statistics.median on the lengths of the responses # TODO 3: return (top_three, median_length) ... top, median_len = report(["yes", "no", "yes", "maybe", "yes", "no", "absolutely"]) print("Top 3 :", ", ".join(f"{word} ({n})" for word, n in top)) print("Median len :", median_len)
Hint 1 — Counter already does the whole first half
Counter(responses) builds the frequency table in one call, and
.most_common(3) returns the top three as a list of
(item, count) pairs, already sorted. That replaces a dict, a loop
and a sort.
Hint 2 — statistics has more than mean
statistics.median takes any sequence of numbers, so build the list
of lengths first: [len(r) for r in responses]. Median is the middle
value — for an even count it averages the middle two, which is why it can return
a float.
Show full solution
from collections import Counter import statistics def report(responses): top_three = Counter(responses).most_common(3) median_length = statistics.median(len(r) for r in responses) return top_three, median_length top, median_len = report(["yes", "no", "yes", "maybe", "yes", "no", "absolutely"]) print("Top 3 :", ", ".join(f"{word} ({n})" for word, n in top)) print("Median len :", median_len) # Top 3 : yes (3), no (2), maybe (1) # Median len : 3
Two lines of real work, both from modules that ship with Python. Written by hand
this is a dict, a loop, a sorted with a key function, a manual sort of the
lengths and an index calculation that is easy to get wrong for even-length
lists — perhaps twenty lines, with the median off by one in a way you would not
notice for months. Before installing anything, it is worth a minute checking
whether the standard library already has it.
Recap
- The standard library ships with Python — no install. Most beginner problems are already solved here.
- pip installs third-party packages from PyPI.
pip install <name>and you canimport <name>. - requirements.txt +
pip freezesaves your project's dependencies. - venv isolates your project's packages. Use it always.
- When in doubt, search before writing — there's probably a module.
This concludes the Fundamentals path. You now have everything you need to write small, useful, real Python programs that read files, handle errors, work with dates, hash things, fetch URLs, and persist data. Time to start building projects — or to dive into the Intermediate path for OOP, decorators, comprehensions, and the techniques that take you from beginner to engineer.
Sources: adapted from Python official documentation (Standard Library overview) and the Python Packaging User Guide. PSF License.