Project: Image Batch Resizer
1 · The lesson
readYou'll build a CLI tool that takes a directory full of photos and spits out web-sized copies — preserving aspect ratios, skipping corrupt files, and running in parallel. The kind of utility that saves you twenty minutes every time you upload a holiday album.
What you'll practice: third-party packages (Pillow), pathlib filesystem iteration, per-file error isolation, argparse, ThreadPoolExecutor, idempotent jobs.
Step 1 — The Bare-Bones Version
Install Pillow first — it's the canonical Python imaging library, a maintained fork of the original PIL.
pip install Pillow
Then the simplest possible resize:
from PIL import Image img = Image.open("photo.jpg") img = img.resize((800, 600)) img.save("photo_small.jpg") print("done")
Three lines of real work. Run it on any photo.jpg and you'll get photo_small.jpg next to it. Pillow handles the JPEG decode, the resampling, and the encode — all you supplied was the new dimensions.
Step 2 — Preserve the Aspect Ratio
resize((800, 600)) is a trap. If your source is 4000x3000, you'll get an 800x600 — that happens to match because the ratio is the same. But on a 4000x2250 source, resize((800, 600)) stretches it vertically. Squashed faces. Sad customers.
Use thumbnail() instead:
from PIL import Image img = Image.open("photo.jpg") img.thumbnail((800, 600)) # in-place, preserves aspect ratio img.save("photo_small.jpg") print(img.size) # e.g. (800, 450) — never bigger than the box
The difference:
| Method | Behaviour |
|---|---|
resize((w, h)) | Forces exact w x h. Can stretch. |
thumbnail((w, h)) | In-place. Scales DOWN to fit inside (w, h), keeping ratio. Never enlarges. |
thumbnail() is what you want for "make this fit in a 1200px box". resize() is for when you genuinely need exact pixels (thumbnails on a grid, ML model inputs).
Step 3 — Batch Over a Directory
One photo is a script. A thousand photos is a tool. Use pathlib to walk the folder and a context manager so each file handle closes cleanly.
from pathlib import Path from PIL import Image source = Path("photos") destination = Path("photos_resized") destination.mkdir(exist_ok=True) # idempotent: no error if it already exists for path in source.glob("*.jpg"): with Image.open(path) as img: img.thumbnail((1200, 1200)) out = destination / path.name img.save(out) print(f" {path.name} → {out}")
A few notes:
glob("*.jpg")returns only.jpgfiles. For JPEG + PNG + WebP, useglob("*.[jJpPwW]*")or iterate and filter on.suffix.lower().with Image.open(path) as img:ensures the file is closed even ifsave()raises. Pillow'sopen()is lazy — it doesn't read pixel data until you ask, so the context manager matters for cleanup of file descriptors on Windows in particular.mkdir(exist_ok=True)makes the script safe to re-run. NoFileExistsErroron the second run.
Step 4 — Handle Errors Per File
One corrupt JPEG in a folder of 500 should NOT abort the batch. Wrap each file in try/except — see exceptions — and keep a tally.
from pathlib import Path from PIL import Image, UnidentifiedImageError source = Path("photos") destination = Path("photos_resized") destination.mkdir(exist_ok=True) ok, failed = 0, [] for path in source.glob("*.jpg"): try: with Image.open(path) as img: img.thumbnail((1200, 1200)) img.save(destination / path.name) ok += 1 except (UnidentifiedImageError, OSError) as e: # UnidentifiedImageError: not actually an image # OSError: truncated file, permission denied, disk full failed.append((path.name, str(e))) print(f"{ok} resized, {len(failed)} failed") for name, err in failed: print(f" ⚠️ {name}: {err}")
Catch the specific exceptions you can handle — UnidentifiedImageError for "not an image", OSError for I/O problems. A bare except: would also swallow KeyboardInterrupt, which means Ctrl-C wouldn't stop the script. Always be specific.
The pattern of "collect failures, report at the end" beats "die on the first error" for any batch tool.
Step 5 — Real CLI Arguments
Hardcoded paths are fine while you're learning, but the tool earns its keep when you can point it at any folder. Use argparse.
import argparse from pathlib import Path from PIL import Image, UnidentifiedImageError def parse_args(): p = argparse.ArgumentParser(description="Batch-resize images.") p.add_argument("input_dir", type=Path, help="folder containing source images") p.add_argument("output_dir", type=Path, help="folder to write resized images") p.add_argument("--max-width", type=int, default=1200, help="longest side in pixels (default 1200)") p.add_argument("--format", choices=["jpeg", "png", "webp"], default=None, help="convert all outputs to this format (default: keep original)") p.add_argument("--quality", type=int, default=85, help="JPEG/WebP quality 1-95 (default 85)") p.add_argument("--recursive", action="store_true", help="include sub-directories") return p.parse_args() def find_images(root, recursive): pattern = "**/*" if recursive else "*" exts = {".jpg", ".jpeg", ".png", ".webp", ".bmp", ".tiff"} for p in root.glob(pattern): if p.is_file() and p.suffix.lower() in exts: yield p def main(): args = parse_args() args.output_dir.mkdir(parents=True, exist_ok=True) for src in find_images(args.input_dir, args.recursive): try: with Image.open(src) as img: img.thumbnail((args.max_width, args.max_width)) ext = f".{args.format}" if args.format else src.suffix dst = args.output_dir / (src.stem + ext) img.save(dst, quality=args.quality) print(f" {src.name} → {dst.name}") except (UnidentifiedImageError, OSError) as e: print(f" ⚠️ {src.name}: {e}") # In a real script: # if __name__ == "__main__": # main()
Usage:
python resize.py ./photos ./web --max-width 1200 --format webp --quality 80 --recursive
argparse gives you --help for free, type-checks the int arguments, and refuses unknown flags. choices= enforces the allowed format list — typo --format webm and argparse rejects it before your code runs.
Step 6 — Polished Final: Parallel, Idempotent, Progress
A 500-photo batch is I/O- and CPU-bound. Pillow releases the GIL during the heavy resampling, so threads (not processes — see multiprocessing for when you'd switch) give a real speedup with almost no code. Add a progress indicator, skip files that already exist, and you have a tool you'd actually pin to your dock.
import argparse from concurrent.futures import ThreadPoolExecutor, as_completed from pathlib import Path from PIL import Image, UnidentifiedImageError IMAGE_EXTS = {".jpg", ".jpeg", ".png", ".webp", ".bmp", ".tiff"} def find_images(root, recursive): pattern = "**/*" if recursive else "*" for p in root.glob(pattern): if p.is_file() and p.suffix.lower() in IMAGE_EXTS: yield p def resize_one(src, output_dir, max_dim, fmt, quality, skip_existing): ext = f".{fmt}" if fmt else src.suffix dst = output_dir / (src.stem + ext) if skip_existing and dst.exists(): return ("skipped", src.name) try: with Image.open(src) as img: img.thumbnail((max_dim, max_dim)) save_kwargs = {} if (fmt or src.suffix.lower()) in (".jpg", ".jpeg", "jpeg", "webp", ".webp"): save_kwargs["quality"] = quality img.save(dst, **save_kwargs) return ("ok", src.name) except (UnidentifiedImageError, OSError) as e: return ("failed", f"{src.name}: {e}") def main(): p = argparse.ArgumentParser(description="Batch-resize images in parallel.") p.add_argument("input_dir", type=Path) p.add_argument("output_dir", type=Path) p.add_argument("--max-width", type=int, default=1200) p.add_argument("--format", choices=["jpeg", "png", "webp"], default=None) p.add_argument("--quality", type=int, default=85) p.add_argument("--recursive", action="store_true") p.add_argument("--workers", type=int, default=8) p.add_argument("--force", action="store_true", help="re-resize even if output exists") args = p.parse_args([]) # demo: parse no args; in real use: p.parse_args() args.input_dir = Path("photos") args.output_dir = Path("photos_resized") args.output_dir.mkdir(parents=True, exist_ok=True) images = list(find_images(args.input_dir, args.recursive)) total = len(images) print(f"Found {total} images. Resizing with {args.workers} workers...") counts = {"ok": 0, "skipped": 0, "failed": 0} with ThreadPoolExecutor(max_workers=args.workers) as pool: futures = [ pool.submit(resize_one, src, args.output_dir, args.max_width, args.format, args.quality, not args.force) for src in images ] for i, fut in enumerate(as_completed(futures), 1): status, msg = fut.result() counts[status] += 1 print(f"\r {i}/{total} ok={counts['ok']} skipped={counts['skipped']} failed={counts['failed']}", end="") print() # newline after the progress line if counts["failed"]: print(f"⚠️ {counts['failed']} failures — check the log.") # In a real script: if __name__ == "__main__": main()
Three quality-of-life wins layered on:
- Threads —
ThreadPoolExecutorparallelises the I/O and the (GIL-releasing) resampling. On an 8-core laptop, a 500-photo batch drops from ~90s to ~15s. Notetqdmis the nicer progress bar option —pip install tqdmand wrapas_completed(futures)intqdm(...). - Idempotency —
skip_existingmeans re-running the script after adding a few photos only processes the new ones. The script becomes a cache. - Format conversion —
--format webprecodes everything. WebP is ~30% smaller than JPEG at the same visual quality. Free bandwidth.
Stretch Goals
1. EXIF preservation — img.info.get("exif") then pass exif= to save(). Otherwise the resized copy loses rotation, camera, GPS metadata.
2. Auto-orient before resize — from PIL import ImageOps; img = ImageOps.exif_transpose(img). Many phone photos are stored sideways with an EXIF tag saying "actually rotate me 90°"; if you resize without applying the orientation, your "landscape" thumbnail ends up portrait.
3. Watermark overlay — ImageDraw.text() or Image.paste() a logo onto the bottom-right corner before saving.
4. Smart crop for variants — generate both a 1200x1200 landscape and a 1080x1920 portrait from each source, centred on the largest face (uses opencv-python for face detection).
5. --replace mode — after a successful resize, delete the original. Dangerous; gate behind an explicit confirmation prompt and an --i-mean-it flag.
🎯 Your Turn — compute_new_size
thumbnail() does this internally, but knowing the maths is useful when you're rolling your own resampler, drawing CSS layouts, or sizing canvas elements. Given an image's original dimensions and a maximum dimension (longest side), return the new (width, height) that fits the constraint while preserving aspect ratio. Never enlarge — if the image is already smaller than max_dim, return it unchanged.
def compute_new_size(orig_w, orig_h, max_dim): """Scale (orig_w, orig_h) so the longest side equals max_dim, preserving aspect ratio. Never upscale. Return ints. """ # TODO 1: if both sides already fit, return (orig_w, orig_h) # TODO 2: compute the scale factor: max_dim / longest_side # TODO 3: apply scale to both dimensions, round to int pass # Test assert compute_new_size(4000, 3000, 1200) == (1200, 900) assert compute_new_size(3000, 4000, 1200) == (900, 1200) assert compute_new_size(800, 600, 1200) == (800, 600) # no upscale assert compute_new_size(1200, 1200, 1200) == (1200, 1200) # exact fit print("all passed")
Hint 1 — Which side is longest?
longest = max(orig_w, orig_h). The scale factor is then max_dim / longest.
Hint 2 — Don't upscale
Iflongest <= max_dim, just return the original dimensions. Otherwise compute the scale and apply.
Show full solution
def compute_new_size(orig_w, orig_h, max_dim): longest = max(orig_w, orig_h) if longest <= max_dim: return (orig_w, orig_h) scale = max_dim / longest return (round(orig_w * scale), round(orig_h * scale)) # Verify for w, h, m, expected in [ (4000, 3000, 1200, (1200, 900)), (3000, 4000, 1200, (900, 1200)), (800, 600, 1200, (800, 600)), (1200, 1200, 1200, (1200, 1200)), (5000, 1000, 1000, (1000, 200)), ]: got = compute_new_size(w, h, m) assert got == expected, f"({w},{h}) max={m}: got {got}, want {expected}" print(f" ({w}, {h}) max={m} → {got}")
The "never upscale" rule matters — enlarging blurs without adding detail. Pillow's thumbnail() follows the same rule.
What You Learned
- Pillow basics —
Image.open(),resize()vsthumbnail(),save()with quality/format options. pathlib.Path.glob()for filtering files by extension.- Per-file error isolation — one bad file doesn't kill the batch.
argparsefor real CLI tools with--helpand type-checked arguments.ThreadPoolExecutorfor parallel I/O-bound work — works well here because Pillow releases the GIL.- Idempotent jobs — skip-if-exists makes scripts safely re-runnable.
You've built a tool you can drop into any project — blog post images, ML dataset prep, customer photo uploads. The skip-if-exists pattern in particular shows up everywhere: build systems, data pipelines, CDN caches. You now know what they're doing under the hood.
Next: Rule-Based Chatbot — pattern-matching, conversation state, and the limits of "if-statement AI".