PythonMastery
intermediate 18 min read · lesson 5 of 12 in Python How-To

Task Automation: Scripting the Boring Parts

1 · The lesson

read

The rule is simple: if you've done it manually three times, automate it. The fourth time is wasted effort; the fifth time you'll forget a step and break something. Python is the lingua franca of system automation because the stdlib already wraps every primitive you need — file ops, process control, scheduling, locking — and the third-party gaps are filled by two or three well-known libraries.

This lesson is the toolbox: subprocess for shelling out, sched/cron/Task Scheduler for time-based triggers, watchdog for filesystem events, file locking to keep jobs from colliding, and the idempotency mindset that turns a fragile cron into a system you can actually retry.


1. The Automation Mindset

Three categories cover almost every automation job you'll write:

  • Batch file ops — rename a thousand camera-roll files, dedupe a directory tree, archive last quarter's logs. Pure pathlib work — see fileio for the streaming patterns.
  • Scheduled jobs — generate a report every morning, back up a database every hour, email a digest every Monday. Triggered by time.
  • System glue — call ffmpeg, pg_dump, git, or any other CLI from Python, parse its output, and feed the result somewhere else. Python becomes the orchestrator.

The third category is where subprocess earns its keep.


2. subprocess.run — The One You Should Use

subprocess.run is the modern, blocking, "do this and tell me how it went" API. The full incantation:

python
import subprocess

result = subprocess.run(
    ["git", "status", "--short"],
    capture_output=True,
    text=True,
    check=True,
)

print(result.stdout)        # the captured stdout as a str
print(result.stderr)        # likewise stderr
print(result.returncode)    # 0 on success (anything else raised CalledProcessError)

Four flags that matter:

  • capture_output=True — grab stdout and stderr into the result object instead of letting them stream to the terminal.
  • text=True — decode the captured bytes as text. Without it you get bytes and have to .decode() everywhere. Always pass it for human-readable commands.
  • check=True — raise subprocess.CalledProcessError on a non-zero exit code. Without it, you have to inspect returncode yourself and remember to handle failure. With it, failures look like every other Python exception.
  • cwd="/some/path" and env={...} — run in a different working directory or with a custom environment. Both are common in build/deploy scripts.

If you only remember one signature, remember that one.


3. shell=True — The Footgun

python
# DANGER — string is interpreted by the shell
subprocess.run(f"ls {user_supplied_path}", shell=True)
+ setup added so this can run · defines subprocess, user_supplied_path
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

subprocess = _AutoMock('subprocess')
user_supplied_path = _AutoMock('user_supplied_path')

If user_supplied_path is "; rm -rf ~", the shell happily runs both commands. This is shell injection, the same class of bug as SQL injection, and the fix is the same: don't concatenate user input into a command string.

The list form sidesteps the shell entirely:

python
# SAFE — arguments are passed directly to the executable, no shell parsing
subprocess.run(["ls", user_supplied_path])
+ setup added so this can run · defines subprocess, user_supplied_path
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

subprocess = _AutoMock('subprocess')
user_supplied_path = _AutoMock('user_supplied_path')

Reach for shell=True only when you genuinely need shell features — pipes, globs, &&, environment variable expansion — and even then, never with untrusted input. 95% of subprocess.run calls should be the list form.


4. Streaming Output for Long-Running Commands

subprocess.run is blocking — it waits for the child to finish before you see anything. For a build that takes ten minutes, that's painful. subprocess.Popen gives you a live handle:

python
import subprocess

proc = subprocess.Popen(
    ["pytest", "-v"],
    stdout=subprocess.PIPE,
    stderr=subprocess.STDOUT,         # merge stderr into stdout
    text=True,
    bufsize=1,                        # line-buffered
)

for line in proc.stdout:
    print(line, end="")               # stream test output as it happens
    # parse, filter, or react here

proc.wait()
if proc.returncode != 0:
    raise RuntimeError(f"pytest failed with exit {proc.returncode}")

bufsize=1 plus iterating proc.stdout gives you a line-at-a-time stream. You can grep, tail, or trigger side-effects as output arrives — useful for CI logs, deploy scripts, and any "show progress" use case.


5. Batch File Ops — pathlib + Glob

The bread-and-butter automation. Three patterns you'll write over and over:

python
from pathlib import Path
from datetime import datetime

# Rename — flatten "IMG_1234.jpg" to "2026-05-14_1234.jpg"
for img in Path("photos").glob("IMG_*.jpg"):
    stamp = datetime.fromtimestamp(img.stat().st_mtime).strftime("%Y-%m-%d")
    img.rename(img.parent / f"{stamp}_{img.stem[4:]}.jpg")

# Dedupe by content hash — keep one copy, delete the rest
import hashlib
seen = {}
for f in Path("downloads").rglob("*"):
    if not f.is_file():
        continue
    digest = hashlib.sha256(f.read_bytes()).hexdigest()
    if digest in seen:
        f.unlink()                    # duplicate — delete
    else:
        seen[digest] = f

# Archive — move anything older than 30 days into archive/
import time
cutoff = time.time() - 30 * 86400
archive = Path("archive")
archive.mkdir(exist_ok=True)
for f in Path("logs").glob("*.log"):
    if f.stat().st_mtime < cutoff:
        f.rename(archive / f.name)

For huge files in a dedupe pass, hash in chunks instead of read_bytes() — see fileio Section 6. The same pathlib vocabulary scales from "rename five files" to "process a million-file dataset."


6. Scheduling — Pick Your Tier

Four tiers of scheduling, in increasing order of weight:

In-process — sched.scheduler

For "run this every N seconds while the program is alive." Lightweight, stdlib, no external dependencies:

python
import sched, time

s = sched.scheduler(time.time, time.sleep)

def heartbeat():
    print("ping", time.strftime("%H:%M:%S"))
    s.enter(5, 1, heartbeat)          # re-schedule itself

s.enter(0, 1, heartbeat)
s.run()                               # blocks; runs scheduled events in order

sched is fine for short-lived daemons and one-off scripts. Don't build a year-long job scheduler on it — your process will be restarted by something eventually.

OS-level — cron / Task Scheduler

For "run this every day at 03:00" on a server, write a normal Python script and let the OS handle the trigger:

python
# crontab -e on Linux/macOS
0 3 * * * /usr/bin/python3 /opt/app/daily_report.py >> /var/log/daily_report.log 2>&1

On Windows, Task Scheduler does the same job through a GUI or schtasks.exe. This is the default for production: the OS owns retries, logs, and survival across reboots; your Python is just the payload.

Library — APScheduler

When you need cron expressions inside a long-running Python process (a web app, a worker) and want persistence (jobs survive restarts), APScheduler is the answer. Supports interval, cron, and date triggers; backs jobs by SQLAlchemy, Redis, or MongoDB.

Distributed — Celery beat, Airflow, Prefect

For web apps with worker fleets, Celery beat is the standard cron-replacement that fans jobs out to workers. For data pipelines with dependencies between tasks, Airflow or Prefect. Flagged in passing — these are weeks of learning, not minutes.


7. File Watching — watchdog

watchdog reacts to filesystem events instead of polling. Drop a file into incoming/ and your script processes it within milliseconds, without a 10-second poll loop burning CPU:

python
# pip install watchdog
from watchdog.observers import Observer
from watchdog.events import FileSystemEventHandler

class IncomingHandler(FileSystemEventHandler):
    def on_created(self, event):
        if event.is_directory:
            return
        print(f"new file: {event.src_path}")
        # process(event.src_path)

observer = Observer()
observer.schedule(IncomingHandler(), path="incoming", recursive=False)
observer.start()
try:
    while True:
        time.sleep(1)
finally:
    observer.stop()
    observer.join()
+ setup added so this can run · defines time
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

time = _AutoMock('time')

Use this for "process every file as it lands" workflows — uploads, scanner outputs, exported reports. The alternative — polling os.listdir every few seconds — works but wastes cycles and adds latency.


8. HTTP Polling vs Webhooks

When you're integrating with another system, the question is who initiates contact:

ApproachYou doWhen to pick
PollingCall their API every N minutes asking "anything new?"The other side has no webhook support; you control the cadence.
WebhooksExpose an HTTP endpoint; they POST when something happens.Low-latency reaction needed; the API offers them; you can host a public endpoint.

Polling is simpler to start with and harder to scale (latency floor = poll interval; wasted requests when nothing changes). Webhooks are near-instant but require a public URL, signature verification, and idempotent handling because the same event may arrive twice.


9. Idempotency — Design for Retries

An automation that can be safely re-run is an automation you can retry. One that can't is one that breaks the moment cron fires twice or you re-run a failed batch.

python
# NON-IDEMPOTENT — re-running creates duplicate rows
def import_orders(rows):
    for row in rows:
        db.insert(row)

# IDEMPOTENT — UPSERT keyed by order_id; re-runs produce the same final state
def import_orders(rows):
    for row in rows:
        db.upsert(row, key="order_id")
+ setup added so this can run · defines db
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

db = _AutoMock('db')

Three idempotency techniques you'll reach for:

  • UPSERT instead of INSERT — same key, same result on re-run.
  • Check-then-act on a marker — write a done.flag file at the end; on retry, skip if it exists.
  • Compute-then-rename — generate report.csv.tmp, then atomically rename to report.csv. Partial work doesn't pollute the final state (see fileio Section 7 for the pattern).

Cron will fire your job at 03:00 sharp whether yesterday's run finished or not. Design accordingly.


10. Locking — Stopping Overlapping Runs

Two cron invocations of the same script can overlap if the first runs long. The cleanest cross-platform fix is the filelock library:

python
# pip install filelock
from filelock import FileLock, Timeout

lock = FileLock("/tmp/daily_backup.lock")
try:
    with lock.acquire(timeout=1):
        run_backup()
except Timeout:
    print("another backup is already running; skipping")
+ setup added so this can run · defines run_backup
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def run_backup(*_a, **_kw):
    print('-> run_backup() called')
    return _AutoMock('run_backup()')

Stdlib alternatives exist but are platform-specific: fcntl.flock on Unix, msvcrt.locking on Windows. filelock wraps both. The pattern is the same everywhere — take the lock, do the work, release on exit, refuse to start if you couldn't get it.


11. The try / except / log / notify Pattern

The full skeleton of a maintainable automation:

python
import logging
import smtplib

logger = logging.getLogger(__name__)

def daily_job():
    try:
        do_the_work()
    except Exception:
        logger.exception("daily_job failed")
        notify_oncall("daily_job failed — see logs")
        raise                             # let the runner record non-zero exit
+ setup added so this can run · defines do_the_work, notify_oncall
# Lightweight mock for objects whose attributes/methods aren't critical
class _AutoMock:
    def __init__(self, name='mock'): self._name = name
    def __getattr__(self, k): return _AutoMock(self._name + '.' + k)
    def __call__(self, *a, **kw):
        print('-> ' + self._name + '() called')
        return _AutoMock(self._name + '()')
    def __repr__(self): return '<mock ' + self._name + '>'
    def __str__(self): return '<mock ' + self._name + '>'
    def __bool__(self): return True
    def __iter__(self): return iter([])
    def __len__(self): return 0
    def __getitem__(self, k): return _AutoMock(self._name + '[...]')
    def __setitem__(self, k, v): pass
    def __enter__(self): return self
    def __exit__(self, *a): return False
    async def __aenter__(self): return self
    async def __aexit__(self, *a): return False
    def __add__(self, o): return self
    def __radd__(self, o): return self
    def __sub__(self, o): return self
    def __mul__(self, o): return self
    def __rmul__(self, o): return self
    def __truediv__(self, o): return self
    def __eq__(self, o): return isinstance(o, _AutoMock)
    def __hash__(self): return hash(self._name)
    def __lt__(self, o): return True
    def __le__(self, o): return True
    def __gt__(self, o): return False
    def __ge__(self, o): return False
    def __mro_entries__(self, bases): return (object,)

def do_the_work(*_a, **_kw):
    print('-> do_the_work() called')
    return _AutoMock('do_the_work()')
def notify_oncall(*_a, **_kw):
    print('-> notify_oncall() called')
    return _AutoMock('notify_oncall()')

Three things happen on failure: a full traceback hits the logs (logger.exception includes it — see logging), the on-call human gets paged, and the process exits non-zero so cron/systemd marks it as failed. Without the notify step, cron mails the failure to a root account nobody reads and you find out three days later.

Wide try/except should still be narrow about what it pretends to know — log everything, then re-raise. Compare with exceptions Section 10.


12. Common Mistakes

1. shell=True with user input
Shell injection. Use the list form.

2. Not capturing stderr
If check=False and you only capture stdout, a failing command looks identical to a silent one. Capture both, or pipe stderr=subprocess.STDOUT to merge them.

3. Long-running scripts with no progress logs
Half an hour in, is it still working or stuck on a hung connection? Log every loop iteration, every external call, every state change. Cheap insurance.

4. Overlapping cron jobs
Job at 03:00 takes 90 minutes once, runs into the 04:00 invocation, both compete on the same files, both corrupt the output. Use a lock (see Section 10).

5. Hardcoded paths

python
# BAD — breaks the moment the script moves
config = open("/home/surya/app/config.json").read()

# GOOD — relative to the script itself
from pathlib import Path
HERE = Path(__file__).resolve().parent
config = (HERE / "config.json").read_text(encoding="utf-8")

6. Wide try/except Exception: pass
Hides every bug. Catch what you can handle; log everything else; re-raise. See exceptions.

7. No failure notification
Cron silently emails root@localhost. Nobody reads it. Wire failures into the same Slack/PagerDuty/email channel you actually monitor.


🎯 Your Turn — Daily Backup Script

Write daily_backup(source_dir, archive_dir) that:

1. Creates a timestamped .tar.gz archive of source_dir inside archive_dir/.
2. Deletes any archive in archive_dir/ older than 30 days.
3. Logs each step at INFO level; logs failures with traceback at ERROR.

The solution should show both approaches — the subprocess form (calling tar) and the pure-stdlib tarfile form.

python
from pathlib import Path

def daily_backup(source_dir, archive_dir):
    # TODO 1: ensure archive_dir exists
    # TODO 2: build a timestamped archive filename
    # TODO 3: create the .tar.gz of source_dir
    # TODO 4: delete archives older than 30 days
    # TODO 5: log each step; on error, log with traceback and re-raise
    ...
Hint 1 — Timestamped filename Use datetime.now(tz=timezone.utc).strftime("%Y%m%d_%H%M%S") and embed it in the filename: backup_20260514_093000.tar.gz. ISO-style timestamps sort lexicographically.
Hint 2 — Pruning by age For each file in archive_dir, check f.stat().st_mtime against time.time() - 30 * 86400. If older, f.unlink().
Show full solution
python
import logging
import subprocess
import tarfile
import time
from datetime import datetime, timezone
from pathlib import Path

logger = logging.getLogger(__name__)


def daily_backup(source_dir, archive_dir, *, use_subprocess=False):
    """Create a timestamped tar.gz of source_dir; prune archives > 30 days old."""
    source_dir = Path(source_dir).resolve()
    archive_dir = Path(archive_dir).resolve()
    archive_dir.mkdir(parents=True, exist_ok=True)

    stamp = datetime.now(tz=timezone.utc).strftime("%Y%m%d_%H%M%S")
    out_path = archive_dir / f"backup_{stamp}.tar.gz"

    try:
        logger.info("creating archive %s from %s", out_path, source_dir)

        if use_subprocess:
            # Approach A — shell out to tar; useful when the system tar
            # has flags or compression you want and you trust the binary.
            subprocess.run(
                ["tar", "-czf", str(out_path), "-C", str(source_dir.parent), source_dir.name],
                check=True,
                capture_output=True,
                text=True,
            )
        else:
            # Approach B — stdlib tarfile; no external dependency, portable to Windows.
            with tarfile.open(out_path, "w:gz") as tar:
                tar.add(source_dir, arcname=source_dir.name)

        logger.info("archive created (%.1f KB)", out_path.stat().st_size / 1024)

        # Prune old archives — anything older than 30 days
        cutoff = time.time() - 30 * 86400
        pruned = 0
        for f in archive_dir.glob("backup_*.tar.gz"):
            if f.stat().st_mtime < cutoff:
                logger.info("pruning old archive %s", f.name)
                f.unlink()
                pruned += 1
        logger.info("done — pruned %d old archive(s)", pruned)

    except subprocess.CalledProcessError as e:
        logger.exception("tar failed: stderr=%s", e.stderr)
        raise
    except Exception:
        logger.exception("daily_backup failed")
        raise


if __name__ == "__main__":
    logging.basicConfig(
        level=logging.INFO,
        format="%(asctime)s %(levelname)s %(name)s: %(message)s",
    )
    daily_backup("./data", "./archive")

What this does:

  • Resolves paths early — Path.resolve() makes the script work whether you pass it a relative or absolute path; no surprises if cwd changes.
  • UTC timestamp in the filename — same string sorts chronologically and never has a DST ambiguity (see datetime).
  • Two implementations — subprocess for systems with a robust tar, tarfile for portability (works on Windows, no external dep). The function picks via a keyword argument.
  • logger.exception inside the failure branch — captures the full traceback automatically.
  • check=True + capture_output=True — non-zero exit raises, and the stderr is in the log on failure rather than lost to stderr.
  • Glob filter backup_*.tar.gz — the prune step only touches files this script created, never random files a user dropped in archive_dir.

For production-grade hardening you'd add: a file lock (see Section 10) so two backups can't run at once; atomic write (write to .tmp and rename, see fileio Section 7) so a half-written archive isn't kept; a webhook/email notifier on failure. All of those are 5-10 extra lines on top of this skeleton.


What You Learned

  • subprocess.run(["cmd", "arg"], capture_output=True, text=True, check=True) is the modern default. List args, not strings; check=True for fail-loud.
  • shell=True with user input is shell injection. Use the list form.
  • Popen + iterate stdout for live streaming of long-running commands.
  • Scheduling tiers: sched in-process, cron/Task Scheduler at OS level, APScheduler for cron expressions inside a daemon, Celery/Airflow for distributed.
  • watchdog for filesystem-event-driven automation instead of polling.
  • Idempotency — design jobs so re-running is safe; UPSERT, markers, atomic rename.
  • File locks (filelock library) to stop overlapping cron runs.
  • try/except → logger.exception → notify → re-raise is the maintainable failure pattern.
  • Path(__file__).resolve().parent instead of hardcoded paths.

Next: Logging & Debugging — because every automation script you write is one production incident away from needing real logs.