PythonMastery
intermediate 30 min read · lesson 14 of 15 in Projects

Project: Image Batch Resizer

1 · The lesson

read

You'll build a CLI tool that takes a directory full of photos and spits out web-sized copies — preserving aspect ratios, skipping corrupt files, and running in parallel. The kind of utility that saves you twenty minutes every time you upload a holiday album.

What you'll practice: third-party packages (Pillow), pathlib filesystem iteration, per-file error isolation, argparse, ThreadPoolExecutor, idempotent jobs.


Step 1 — The Bare-Bones Version

Install Pillow first — it's the canonical Python imaging library, a maintained fork of the original PIL.

python
pip install Pillow

Then the simplest possible resize:

python
from PIL import Image

img = Image.open("photo.jpg")
img = img.resize((800, 600))
img.save("photo_small.jpg")
print("done")

Three lines of real work. Run it on any photo.jpg and you'll get photo_small.jpg next to it. Pillow handles the JPEG decode, the resampling, and the encode — all you supplied was the new dimensions.


Step 2 — Preserve the Aspect Ratio

resize((800, 600)) is a trap. If your source is 4000x3000, you'll get an 800x600 — that happens to match because the ratio is the same. But on a 4000x2250 source, resize((800, 600)) stretches it vertically. Squashed faces. Sad customers.

Use thumbnail() instead:

python
from PIL import Image

img = Image.open("photo.jpg")
img.thumbnail((800, 600))         # in-place, preserves aspect ratio
img.save("photo_small.jpg")

print(img.size)                   # e.g. (800, 450) — never bigger than the box

The difference:

MethodBehaviour
resize((w, h))Forces exact w x h. Can stretch.
thumbnail((w, h))In-place. Scales DOWN to fit inside (w, h), keeping ratio. Never enlarges.

thumbnail() is what you want for "make this fit in a 1200px box". resize() is for when you genuinely need exact pixels (thumbnails on a grid, ML model inputs).


Step 3 — Batch Over a Directory

One photo is a script. A thousand photos is a tool. Use pathlib to walk the folder and a context manager so each file handle closes cleanly.

python
from pathlib import Path
from PIL import Image

source = Path("photos")
destination = Path("photos_resized")
destination.mkdir(exist_ok=True)        # idempotent: no error if it already exists

for path in source.glob("*.jpg"):
    with Image.open(path) as img:
        img.thumbnail((1200, 1200))
        out = destination / path.name
        img.save(out)
        print(f"  {path.name} → {out}")

A few notes:

  • glob("*.jpg") returns only .jpg files. For JPEG + PNG + WebP, use glob("*.[jJpPwW]*") or iterate and filter on .suffix.lower().
  • with Image.open(path) as img: ensures the file is closed even if save() raises. Pillow's open() is lazy — it doesn't read pixel data until you ask, so the context manager matters for cleanup of file descriptors on Windows in particular.
  • mkdir(exist_ok=True) makes the script safe to re-run. No FileExistsError on the second run.

Step 4 — Handle Errors Per File

One corrupt JPEG in a folder of 500 should NOT abort the batch. Wrap each file in try/except — see exceptions — and keep a tally.

python
from pathlib import Path
from PIL import Image, UnidentifiedImageError

source = Path("photos")
destination = Path("photos_resized")
destination.mkdir(exist_ok=True)

ok, failed = 0, []

for path in source.glob("*.jpg"):
    try:
        with Image.open(path) as img:
            img.thumbnail((1200, 1200))
            img.save(destination / path.name)
        ok += 1
    except (UnidentifiedImageError, OSError) as e:
        # UnidentifiedImageError: not actually an image
        # OSError: truncated file, permission denied, disk full
        failed.append((path.name, str(e)))

print(f"{ok} resized, {len(failed)} failed")
for name, err in failed:
    print(f"  ⚠️  {name}: {err}")

Catch the specific exceptions you can handle — UnidentifiedImageError for "not an image", OSError for I/O problems. A bare except: would also swallow KeyboardInterrupt, which means Ctrl-C wouldn't stop the script. Always be specific.

The pattern of "collect failures, report at the end" beats "die on the first error" for any batch tool.


Step 5 — Real CLI Arguments

Hardcoded paths are fine while you're learning, but the tool earns its keep when you can point it at any folder. Use argparse.

python
import argparse
from pathlib import Path
from PIL import Image, UnidentifiedImageError

def parse_args():
    p = argparse.ArgumentParser(description="Batch-resize images.")
    p.add_argument("input_dir", type=Path, help="folder containing source images")
    p.add_argument("output_dir", type=Path, help="folder to write resized images")
    p.add_argument("--max-width", type=int, default=1200, help="longest side in pixels (default 1200)")
    p.add_argument("--format", choices=["jpeg", "png", "webp"], default=None,
                   help="convert all outputs to this format (default: keep original)")
    p.add_argument("--quality", type=int, default=85, help="JPEG/WebP quality 1-95 (default 85)")
    p.add_argument("--recursive", action="store_true", help="include sub-directories")
    return p.parse_args()

def find_images(root, recursive):
    pattern = "**/*" if recursive else "*"
    exts = {".jpg", ".jpeg", ".png", ".webp", ".bmp", ".tiff"}
    for p in root.glob(pattern):
        if p.is_file() and p.suffix.lower() in exts:
            yield p

def main():
    args = parse_args()
    args.output_dir.mkdir(parents=True, exist_ok=True)

    for src in find_images(args.input_dir, args.recursive):
        try:
            with Image.open(src) as img:
                img.thumbnail((args.max_width, args.max_width))
                ext = f".{args.format}" if args.format else src.suffix
                dst = args.output_dir / (src.stem + ext)
                img.save(dst, quality=args.quality)
                print(f"  {src.name} → {dst.name}")
        except (UnidentifiedImageError, OSError) as e:
            print(f"  ⚠️  {src.name}: {e}")

# In a real script:
# if __name__ == "__main__":
#     main()

Usage:

python
python resize.py ./photos ./web --max-width 1200 --format webp --quality 80 --recursive

argparse gives you --help for free, type-checks the int arguments, and refuses unknown flags. choices= enforces the allowed format list — typo --format webm and argparse rejects it before your code runs.


Step 6 — Polished Final: Parallel, Idempotent, Progress

A 500-photo batch is I/O- and CPU-bound. Pillow releases the GIL during the heavy resampling, so threads (not processes — see multiprocessing for when you'd switch) give a real speedup with almost no code. Add a progress indicator, skip files that already exist, and you have a tool you'd actually pin to your dock.

python
import argparse
from concurrent.futures import ThreadPoolExecutor, as_completed
from pathlib import Path
from PIL import Image, UnidentifiedImageError

IMAGE_EXTS = {".jpg", ".jpeg", ".png", ".webp", ".bmp", ".tiff"}

def find_images(root, recursive):
    pattern = "**/*" if recursive else "*"
    for p in root.glob(pattern):
        if p.is_file() and p.suffix.lower() in IMAGE_EXTS:
            yield p

def resize_one(src, output_dir, max_dim, fmt, quality, skip_existing):
    ext = f".{fmt}" if fmt else src.suffix
    dst = output_dir / (src.stem + ext)

    if skip_existing and dst.exists():
        return ("skipped", src.name)

    try:
        with Image.open(src) as img:
            img.thumbnail((max_dim, max_dim))
            save_kwargs = {}
            if (fmt or src.suffix.lower()) in (".jpg", ".jpeg", "jpeg", "webp", ".webp"):
                save_kwargs["quality"] = quality
            img.save(dst, **save_kwargs)
        return ("ok", src.name)
    except (UnidentifiedImageError, OSError) as e:
        return ("failed", f"{src.name}: {e}")

def main():
    p = argparse.ArgumentParser(description="Batch-resize images in parallel.")
    p.add_argument("input_dir", type=Path)
    p.add_argument("output_dir", type=Path)
    p.add_argument("--max-width", type=int, default=1200)
    p.add_argument("--format", choices=["jpeg", "png", "webp"], default=None)
    p.add_argument("--quality", type=int, default=85)
    p.add_argument("--recursive", action="store_true")
    p.add_argument("--workers", type=int, default=8)
    p.add_argument("--force", action="store_true", help="re-resize even if output exists")
    args = p.parse_args([])    # demo: parse no args; in real use: p.parse_args()

    args.input_dir = Path("photos")
    args.output_dir = Path("photos_resized")
    args.output_dir.mkdir(parents=True, exist_ok=True)

    images = list(find_images(args.input_dir, args.recursive))
    total = len(images)
    print(f"Found {total} images. Resizing with {args.workers} workers...")

    counts = {"ok": 0, "skipped": 0, "failed": 0}

    with ThreadPoolExecutor(max_workers=args.workers) as pool:
        futures = [
            pool.submit(resize_one, src, args.output_dir,
                        args.max_width, args.format, args.quality,
                        not args.force)
            for src in images
        ]
        for i, fut in enumerate(as_completed(futures), 1):
            status, msg = fut.result()
            counts[status] += 1
            print(f"\r  {i}/{total}  ok={counts['ok']} skipped={counts['skipped']} failed={counts['failed']}", end="")

    print()    # newline after the progress line
    if counts["failed"]:
        print(f"⚠️  {counts['failed']} failures — check the log.")

# In a real script: if __name__ == "__main__": main()

Three quality-of-life wins layered on:

  • Threads — ThreadPoolExecutor parallelises the I/O and the (GIL-releasing) resampling. On an 8-core laptop, a 500-photo batch drops from ~90s to ~15s. Note tqdm is the nicer progress bar option — pip install tqdm and wrap as_completed(futures) in tqdm(...).
  • Idempotency — skip_existing means re-running the script after adding a few photos only processes the new ones. The script becomes a cache.
  • Format conversion — --format webp recodes everything. WebP is ~30% smaller than JPEG at the same visual quality. Free bandwidth.

Stretch Goals

1. EXIF preservation — img.info.get("exif") then pass exif= to save(). Otherwise the resized copy loses rotation, camera, GPS metadata.
2. Auto-orient before resize — from PIL import ImageOps; img = ImageOps.exif_transpose(img). Many phone photos are stored sideways with an EXIF tag saying "actually rotate me 90°"; if you resize without applying the orientation, your "landscape" thumbnail ends up portrait.
3. Watermark overlay — ImageDraw.text() or Image.paste() a logo onto the bottom-right corner before saving.
4. Smart crop for variants — generate both a 1200x1200 landscape and a 1080x1920 portrait from each source, centred on the largest face (uses opencv-python for face detection).
5. --replace mode — after a successful resize, delete the original. Dangerous; gate behind an explicit confirmation prompt and an --i-mean-it flag.


🎯 Your Turn — compute_new_size

thumbnail() does this internally, but knowing the maths is useful when you're rolling your own resampler, drawing CSS layouts, or sizing canvas elements. Given an image's original dimensions and a maximum dimension (longest side), return the new (width, height) that fits the constraint while preserving aspect ratio. Never enlarge — if the image is already smaller than max_dim, return it unchanged.

python
def compute_new_size(orig_w, orig_h, max_dim):
    """Scale (orig_w, orig_h) so the longest side equals max_dim,
    preserving aspect ratio. Never upscale. Return ints.
    """
    # TODO 1: if both sides already fit, return (orig_w, orig_h)
    # TODO 2: compute the scale factor: max_dim / longest_side
    # TODO 3: apply scale to both dimensions, round to int
    pass

# Test
assert compute_new_size(4000, 3000, 1200) == (1200, 900)
assert compute_new_size(3000, 4000, 1200) == (900, 1200)
assert compute_new_size(800, 600, 1200) == (800, 600)         # no upscale
assert compute_new_size(1200, 1200, 1200) == (1200, 1200)     # exact fit
print("all passed")
Hint 1 — Which side is longest? longest = max(orig_w, orig_h). The scale factor is then max_dim / longest.
Hint 2 — Don't upscale If longest <= max_dim, just return the original dimensions. Otherwise compute the scale and apply.
Show full solution
python
def compute_new_size(orig_w, orig_h, max_dim):
    longest = max(orig_w, orig_h)
    if longest <= max_dim:
        return (orig_w, orig_h)
    scale = max_dim / longest
    return (round(orig_w * scale), round(orig_h * scale))

# Verify
for w, h, m, expected in [
    (4000, 3000, 1200, (1200, 900)),
    (3000, 4000, 1200, (900, 1200)),
    (800, 600, 1200, (800, 600)),
    (1200, 1200, 1200, (1200, 1200)),
    (5000, 1000, 1000, (1000, 200)),
]:
    got = compute_new_size(w, h, m)
    assert got == expected, f"({w},{h}) max={m}: got {got}, want {expected}"
    print(f"  ({w}, {h}) max={m} → {got}")

The "never upscale" rule matters — enlarging blurs without adding detail. Pillow's thumbnail() follows the same rule.


What You Learned

  • Pillow basics — Image.open(), resize() vs thumbnail(), save() with quality/format options.
  • pathlib.Path.glob() for filtering files by extension.
  • Per-file error isolation — one bad file doesn't kill the batch.
  • argparse for real CLI tools with --help and type-checked arguments.
  • ThreadPoolExecutor for parallel I/O-bound work — works well here because Pillow releases the GIL.
  • Idempotent jobs — skip-if-exists makes scripts safely re-runnable.

You've built a tool you can drop into any project — blog post images, ML dataset prep, customer photo uploads. The skip-if-exists pattern in particular shows up everywhere: build systems, data pipelines, CDN caches. You now know what they're doing under the hood.

Next: Rule-Based Chatbot — pattern-matching, conversation state, and the limits of "if-statement AI".