Task Automation: Scripting the Boring Parts
1 · The lesson
readThe rule is simple: if you've done it manually three times, automate it. The fourth time is wasted effort; the fifth time you'll forget a step and break something. Python is the lingua franca of system automation because the stdlib already wraps every primitive you need — file ops, process control, scheduling, locking — and the third-party gaps are filled by two or three well-known libraries.
This lesson is the toolbox: subprocess for shelling out, sched/cron/Task Scheduler for time-based triggers, watchdog for filesystem events, file locking to keep jobs from colliding, and the idempotency mindset that turns a fragile cron into a system you can actually retry.
1. The Automation Mindset
Three categories cover almost every automation job you'll write:
- Batch file ops — rename a thousand camera-roll files, dedupe a directory tree, archive last quarter's logs. Pure
pathlibwork — see fileio for the streaming patterns. - Scheduled jobs — generate a report every morning, back up a database every hour, email a digest every Monday. Triggered by time.
- System glue — call
ffmpeg,pg_dump,git, or any other CLI from Python, parse its output, and feed the result somewhere else. Python becomes the orchestrator.
The third category is where subprocess earns its keep.
2. subprocess.run — The One You Should Use
subprocess.run is the modern, blocking, "do this and tell me how it went" API. The full incantation:
import subprocess result = subprocess.run( ["git", "status", "--short"], capture_output=True, text=True, check=True, ) print(result.stdout) # the captured stdout as a str print(result.stderr) # likewise stderr print(result.returncode) # 0 on success (anything else raised CalledProcessError)
Four flags that matter:
capture_output=True— grabstdoutandstderrinto the result object instead of letting them stream to the terminal.text=True— decode the captured bytes as text. Without it you getbytesand have to.decode()everywhere. Always pass it for human-readable commands.check=True— raisesubprocess.CalledProcessErroron a non-zero exit code. Without it, you have to inspectreturncodeyourself and remember to handle failure. With it, failures look like every other Python exception.cwd="/some/path"andenv={...}— run in a different working directory or with a custom environment. Both are common in build/deploy scripts.
If you only remember one signature, remember that one.
3. shell=True — The Footgun
# DANGER — string is interpreted by the shell subprocess.run(f"ls {user_supplied_path}", shell=True)
setup added so this can run · defines subprocess, user_supplied_path
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) subprocess = _AutoMock('subprocess') user_supplied_path = _AutoMock('user_supplied_path')
If user_supplied_path is "; rm -rf ~", the shell happily runs both commands. This is shell injection, the same class of bug as SQL injection, and the fix is the same: don't concatenate user input into a command string.
The list form sidesteps the shell entirely:
# SAFE — arguments are passed directly to the executable, no shell parsing subprocess.run(["ls", user_supplied_path])
setup added so this can run · defines subprocess, user_supplied_path
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) subprocess = _AutoMock('subprocess') user_supplied_path = _AutoMock('user_supplied_path')
Reach for shell=True only when you genuinely need shell features — pipes, globs, &&, environment variable expansion — and even then, never with untrusted input. 95% of subprocess.run calls should be the list form.
4. Streaming Output for Long-Running Commands
subprocess.run is blocking — it waits for the child to finish before you see anything. For a build that takes ten minutes, that's painful. subprocess.Popen gives you a live handle:
import subprocess proc = subprocess.Popen( ["pytest", "-v"], stdout=subprocess.PIPE, stderr=subprocess.STDOUT, # merge stderr into stdout text=True, bufsize=1, # line-buffered ) for line in proc.stdout: print(line, end="") # stream test output as it happens # parse, filter, or react here proc.wait() if proc.returncode != 0: raise RuntimeError(f"pytest failed with exit {proc.returncode}")
bufsize=1 plus iterating proc.stdout gives you a line-at-a-time stream. You can grep, tail, or trigger side-effects as output arrives — useful for CI logs, deploy scripts, and any "show progress" use case.
5. Batch File Ops — pathlib + Glob
The bread-and-butter automation. Three patterns you'll write over and over:
from pathlib import Path from datetime import datetime # Rename — flatten "IMG_1234.jpg" to "2026-05-14_1234.jpg" for img in Path("photos").glob("IMG_*.jpg"): stamp = datetime.fromtimestamp(img.stat().st_mtime).strftime("%Y-%m-%d") img.rename(img.parent / f"{stamp}_{img.stem[4:]}.jpg") # Dedupe by content hash — keep one copy, delete the rest import hashlib seen = {} for f in Path("downloads").rglob("*"): if not f.is_file(): continue digest = hashlib.sha256(f.read_bytes()).hexdigest() if digest in seen: f.unlink() # duplicate — delete else: seen[digest] = f # Archive — move anything older than 30 days into archive/ import time cutoff = time.time() - 30 * 86400 archive = Path("archive") archive.mkdir(exist_ok=True) for f in Path("logs").glob("*.log"): if f.stat().st_mtime < cutoff: f.rename(archive / f.name)
For huge files in a dedupe pass, hash in chunks instead of read_bytes() — see fileio Section 6. The same pathlib vocabulary scales from "rename five files" to "process a million-file dataset."
6. Scheduling — Pick Your Tier
Four tiers of scheduling, in increasing order of weight:
In-process — sched.scheduler
For "run this every N seconds while the program is alive." Lightweight, stdlib, no external dependencies:
import sched, time s = sched.scheduler(time.time, time.sleep) def heartbeat(): print("ping", time.strftime("%H:%M:%S")) s.enter(5, 1, heartbeat) # re-schedule itself s.enter(0, 1, heartbeat) s.run() # blocks; runs scheduled events in order
sched is fine for short-lived daemons and one-off scripts. Don't build a year-long job scheduler on it — your process will be restarted by something eventually.
OS-level — cron / Task Scheduler
For "run this every day at 03:00" on a server, write a normal Python script and let the OS handle the trigger:
# crontab -e on Linux/macOS 0 3 * * * /usr/bin/python3 /opt/app/daily_report.py >> /var/log/daily_report.log 2>&1
On Windows, Task Scheduler does the same job through a GUI or schtasks.exe. This is the default for production: the OS owns retries, logs, and survival across reboots; your Python is just the payload.
Library — APScheduler
When you need cron expressions inside a long-running Python process (a web app, a worker) and want persistence (jobs survive restarts), APScheduler is the answer. Supports interval, cron, and date triggers; backs jobs by SQLAlchemy, Redis, or MongoDB.
Distributed — Celery beat, Airflow, Prefect
For web apps with worker fleets, Celery beat is the standard cron-replacement that fans jobs out to workers. For data pipelines with dependencies between tasks, Airflow or Prefect. Flagged in passing — these are weeks of learning, not minutes.
7. File Watching — watchdog
watchdog reacts to filesystem events instead of polling. Drop a file into incoming/ and your script processes it within milliseconds, without a 10-second poll loop burning CPU:
# pip install watchdog from watchdog.observers import Observer from watchdog.events import FileSystemEventHandler class IncomingHandler(FileSystemEventHandler): def on_created(self, event): if event.is_directory: return print(f"new file: {event.src_path}") # process(event.src_path) observer = Observer() observer.schedule(IncomingHandler(), path="incoming", recursive=False) observer.start() try: while True: time.sleep(1) finally: observer.stop() observer.join()
setup added so this can run · defines time
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) time = _AutoMock('time')
Use this for "process every file as it lands" workflows — uploads, scanner outputs, exported reports. The alternative — polling os.listdir every few seconds — works but wastes cycles and adds latency.
8. HTTP Polling vs Webhooks
When you're integrating with another system, the question is who initiates contact:
| Approach | You do | When to pick |
|---|---|---|
| Polling | Call their API every N minutes asking "anything new?" | The other side has no webhook support; you control the cadence. |
| Webhooks | Expose an HTTP endpoint; they POST when something happens. | Low-latency reaction needed; the API offers them; you can host a public endpoint. |
Polling is simpler to start with and harder to scale (latency floor = poll interval; wasted requests when nothing changes). Webhooks are near-instant but require a public URL, signature verification, and idempotent handling because the same event may arrive twice.
9. Idempotency — Design for Retries
An automation that can be safely re-run is an automation you can retry. One that can't is one that breaks the moment cron fires twice or you re-run a failed batch.
# NON-IDEMPOTENT — re-running creates duplicate rows def import_orders(rows): for row in rows: db.insert(row) # IDEMPOTENT — UPSERT keyed by order_id; re-runs produce the same final state def import_orders(rows): for row in rows: db.upsert(row, key="order_id")
setup added so this can run · defines db
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) db = _AutoMock('db')
Three idempotency techniques you'll reach for:
- UPSERT instead of INSERT — same key, same result on re-run.
- Check-then-act on a marker — write a
done.flagfile at the end; on retry, skip if it exists. - Compute-then-rename — generate
report.csv.tmp, then atomically rename toreport.csv. Partial work doesn't pollute the final state (see fileio Section 7 for the pattern).
Cron will fire your job at 03:00 sharp whether yesterday's run finished or not. Design accordingly.
10. Locking — Stopping Overlapping Runs
Two cron invocations of the same script can overlap if the first runs long. The cleanest cross-platform fix is the filelock library:
# pip install filelock from filelock import FileLock, Timeout lock = FileLock("/tmp/daily_backup.lock") try: with lock.acquire(timeout=1): run_backup() except Timeout: print("another backup is already running; skipping")
setup added so this can run · defines run_backup
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def run_backup(*_a, **_kw): print('-> run_backup() called') return _AutoMock('run_backup()')
Stdlib alternatives exist but are platform-specific: fcntl.flock on Unix, msvcrt.locking on Windows. filelock wraps both. The pattern is the same everywhere — take the lock, do the work, release on exit, refuse to start if you couldn't get it.
11. The try / except / log / notify Pattern
The full skeleton of a maintainable automation:
import logging import smtplib logger = logging.getLogger(__name__) def daily_job(): try: do_the_work() except Exception: logger.exception("daily_job failed") notify_oncall("daily_job failed — see logs") raise # let the runner record non-zero exit
setup added so this can run · defines do_the_work, notify_oncall
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) def do_the_work(*_a, **_kw): print('-> do_the_work() called') return _AutoMock('do_the_work()') def notify_oncall(*_a, **_kw): print('-> notify_oncall() called') return _AutoMock('notify_oncall()')
Three things happen on failure: a full traceback hits the logs (logger.exception includes it — see logging), the on-call human gets paged, and the process exits non-zero so cron/systemd marks it as failed. Without the notify step, cron mails the failure to a root account nobody reads and you find out three days later.
Wide try/except should still be narrow about what it pretends to know — log everything, then re-raise. Compare with exceptions Section 10.
12. Common Mistakes
1. shell=True with user input
Shell injection. Use the list form.
2. Not capturing stderr
If check=False and you only capture stdout, a failing command looks identical to a silent one. Capture both, or pipe stderr=subprocess.STDOUT to merge them.
3. Long-running scripts with no progress logs
Half an hour in, is it still working or stuck on a hung connection? Log every loop iteration, every external call, every state change. Cheap insurance.
4. Overlapping cron jobs
Job at 03:00 takes 90 minutes once, runs into the 04:00 invocation, both compete on the same files, both corrupt the output. Use a lock (see Section 10).
5. Hardcoded paths
# BAD — breaks the moment the script moves config = open("/home/surya/app/config.json").read() # GOOD — relative to the script itself from pathlib import Path HERE = Path(__file__).resolve().parent config = (HERE / "config.json").read_text(encoding="utf-8")
6. Wide try/except Exception: pass
Hides every bug. Catch what you can handle; log everything else; re-raise. See exceptions.
7. No failure notification
Cron silently emails root@localhost. Nobody reads it. Wire failures into the same Slack/PagerDuty/email channel you actually monitor.
🎯 Your Turn — Daily Backup Script
Write daily_backup(source_dir, archive_dir) that:
1. Creates a timestamped .tar.gz archive of source_dir inside archive_dir/.
2. Deletes any archive in archive_dir/ older than 30 days.
3. Logs each step at INFO level; logs failures with traceback at ERROR.
The solution should show both approaches — the subprocess form (calling tar) and the pure-stdlib tarfile form.
from pathlib import Path def daily_backup(source_dir, archive_dir): # TODO 1: ensure archive_dir exists # TODO 2: build a timestamped archive filename # TODO 3: create the .tar.gz of source_dir # TODO 4: delete archives older than 30 days # TODO 5: log each step; on error, log with traceback and re-raise ...
Hint 1 — Timestamped filename
Usedatetime.now(tz=timezone.utc).strftime("%Y%m%d_%H%M%S") and embed it in the filename: backup_20260514_093000.tar.gz. ISO-style timestamps sort lexicographically.
Hint 2 — Pruning by age
For each file inarchive_dir, check f.stat().st_mtime against time.time() - 30 * 86400. If older, f.unlink().
Show full solution
import logging import subprocess import tarfile import time from datetime import datetime, timezone from pathlib import Path logger = logging.getLogger(__name__) def daily_backup(source_dir, archive_dir, *, use_subprocess=False): """Create a timestamped tar.gz of source_dir; prune archives > 30 days old.""" source_dir = Path(source_dir).resolve() archive_dir = Path(archive_dir).resolve() archive_dir.mkdir(parents=True, exist_ok=True) stamp = datetime.now(tz=timezone.utc).strftime("%Y%m%d_%H%M%S") out_path = archive_dir / f"backup_{stamp}.tar.gz" try: logger.info("creating archive %s from %s", out_path, source_dir) if use_subprocess: # Approach A — shell out to tar; useful when the system tar # has flags or compression you want and you trust the binary. subprocess.run( ["tar", "-czf", str(out_path), "-C", str(source_dir.parent), source_dir.name], check=True, capture_output=True, text=True, ) else: # Approach B — stdlib tarfile; no external dependency, portable to Windows. with tarfile.open(out_path, "w:gz") as tar: tar.add(source_dir, arcname=source_dir.name) logger.info("archive created (%.1f KB)", out_path.stat().st_size / 1024) # Prune old archives — anything older than 30 days cutoff = time.time() - 30 * 86400 pruned = 0 for f in archive_dir.glob("backup_*.tar.gz"): if f.stat().st_mtime < cutoff: logger.info("pruning old archive %s", f.name) f.unlink() pruned += 1 logger.info("done — pruned %d old archive(s)", pruned) except subprocess.CalledProcessError as e: logger.exception("tar failed: stderr=%s", e.stderr) raise except Exception: logger.exception("daily_backup failed") raise if __name__ == "__main__": logging.basicConfig( level=logging.INFO, format="%(asctime)s %(levelname)s %(name)s: %(message)s", ) daily_backup("./data", "./archive")
What this does:
- Resolves paths early —
Path.resolve()makes the script work whether you pass it a relative or absolute path; no surprises ifcwdchanges. - UTC timestamp in the filename — same string sorts chronologically and never has a DST ambiguity (see datetime).
- Two implementations —
subprocessfor systems with a robusttar,tarfilefor portability (works on Windows, no external dep). The function picks via a keyword argument. logger.exceptioninside the failure branch — captures the full traceback automatically.check=True+capture_output=True— non-zero exit raises, and the stderr is in the log on failure rather than lost to stderr.- Glob filter
backup_*.tar.gz— the prune step only touches files this script created, never random files a user dropped inarchive_dir.
For production-grade hardening you'd add: a file lock (see Section 10) so two backups can't run at once; atomic write (write to .tmp and rename, see fileio Section 7) so a half-written archive isn't kept; a webhook/email notifier on failure. All of those are 5-10 extra lines on top of this skeleton.
What You Learned
subprocess.run(["cmd", "arg"], capture_output=True, text=True, check=True)is the modern default. List args, not strings;check=Truefor fail-loud.shell=Truewith user input is shell injection. Use the list form.Popen+ iteratestdoutfor live streaming of long-running commands.- Scheduling tiers:
schedin-process, cron/Task Scheduler at OS level,APSchedulerfor cron expressions inside a daemon, Celery/Airflow for distributed. watchdogfor filesystem-event-driven automation instead of polling.- Idempotency — design jobs so re-running is safe; UPSERT, markers, atomic rename.
- File locks (
filelocklibrary) to stop overlapping cron runs. try/except→logger.exception→ notify → re-raise is the maintainable failure pattern.Path(__file__).resolve().parentinstead of hardcoded paths.
Next: Logging & Debugging — because every automation script you write is one production incident away from needing real logs.