Environment Management for Python Projects
1 · The lesson
readExamples require Docker / a GitHub repo / a deployment target. Not browser-runnable. Commands shown with the expected output in comments.
"Deploy and pray" is what happens when you have one environment. The same artefact that handled your unit tests yesterday now serves real users; the first time anyone runs it against production-shaped data, traffic, or auth providers is the moment it breaks customers. A grown-up deployment story has at least three environments — dev, staging, prod — and the same code runs in all three with config the only thing that varies.
This lesson is about the shape of multi-environment Python deployment: what changes between environments, what stays the same, how config flows in, how branches and tags map to deploys, and the observability hooks that tell you which environment a bug is from. It builds directly on envconfig, docker, and github-actions.
1. The Three-Environment Baseline
The minimum that doesn't hurt:
| Environment | Where | Who uses it | Data |
|---|---|---|---|
| Development | Your laptop | You | Local, throwaway. Seed scripts and fixtures. |
| Staging | Cloud-hosted, internal-only | Your team, QA, sometimes designers | Synthetic or anonymised production data |
| Production | Cloud-hosted, public | Real users | Real, irreplaceable |
Staging is the environment most teams under-invest in. It mirrors prod's infrastructure shape — same database engine, same versions, same auth provider, same secrets manager, same CDN — but with separate data and separate credentials. Its job is to catch issues prod would catch, before prod sees them. If staging never has bugs, it isn't staging — it's a second prod that you haven't told anyone about.
The single most important property of staging: it must be safe to break. Engineers should feel free to deploy half-finished work, drop tables, restart services. The moment a team starts treating staging like prod (asking for change windows, requiring approvals), it stops doing its job.
2. What Varies Between Environments
The list is shorter than you might think:
- Database URL — separate DB per environment, always.
- External API keys — Stripe test keys in dev/staging, live keys in prod. Same for Anthropic, Twilio, AWS.
- OAuth client IDs / redirect URIs — registered per environment.
- Log level —
DEBUGlocally,INFOin staging,WARNINGorINFOin prod (preference varies). - Feature flags — new features on in staging, off in prod until rollout.
- Rate limits — usually relaxed in staging for load testing, tight in prod.
- Email / SMS providers — sandbox in dev/staging (never spam real users from staging), real in prod.
- CORS allowed origins —
http://localhost:3000in dev, the real domain in prod. - Sentry DSN / log-aggregator endpoint — separate projects per env so prod errors don't get lost in dev noise.
- Domain / base URL — affects email links, webhooks, OAuth redirects.
Everything else — your code, your migrations, your container images, your CI pipeline — is identical. Same Docker image, same git SHA, same dependencies.
3. The 12-Factor Principle (Recap)
From envconfig and worth restating: anything that varies between environments goes in environment variables. Code is identical across environments. This is the whole point of containers + env-var config — the same image runs anywhere, configured by what's in os.environ at startup.
The minimum Python pattern:
import os class Config: DATABASE_URL: str = os.environ["DATABASE_URL"] DEBUG: bool = os.environ.get("DEBUG", "false").lower() == "true" LOG_LEVEL: str = os.environ.get("LOG_LEVEL", "INFO") APP_ENV: str = os.environ.get("APP_ENV", "development")
setup added so this can run · defines
import os # noqa: F401 os.environ.setdefault("DATABASE_URL", "postgresql://user:password@localhost:5432/example") os.environ.setdefault("DEBUG", "false") os.environ.setdefault("LOG_LEVEL", "example-log-level") os.environ.setdefault("APP_ENV", "example-app-env")
For anything beyond ~5 fields, reach for pydantic-settings — typed, validated, automatic .env support, no hand-rolled int()/bool() coercion:
from pydantic_settings import BaseSettings, SettingsConfigDict class Settings(BaseSettings): model_config = SettingsConfigDict(env_file=".env", env_file_encoding="utf-8") app_env: str = "development" database_url: str log_level: str = "INFO" debug: bool = False sentry_dsn: str | None = None feature_new_search: bool = False settings = Settings()
Production preference: pydantic-settings. The validation pays for itself the first time someone sets PORT=eight-thousand and pydantic refuses to start the app with a clear message.
4. The Env-Var Hierarchy
Where does a config value come from? Multiple sources, with a strict precedence:
host / cloud secrets (highest)
↓
shell environment (CI runner, your terminal)
↓
.env file (local dev only — never in prod)
↓
.env.example (committed template, defaults)
↓
code defaults (lowest)In practice:
.env.example— committed. Lists every variable with placeholder values.DATABASE_URL=postgres://user:pass@host/db. Documents what config exists..env— gitignored. Local-dev values. Created by copying.env.exampleand filling in real values.- Shell env / CI secrets — overrides
.env. Useful for one-off runs (DEBUG=true python app.py) and CI matrix variations. - Host secrets — production. Injected by the platform (AWS Secrets Manager, Kubernetes secrets, ECS task definitions). Highest precedence — overrides anything else.
The discipline: no environment-specific files in git. No .env.production, no config.prod.yaml. The closest is docker-compose.prod.yml, which contains references to env vars (${DATABASE_URL}) but never values.
5. Multi-Environment docker compose
Compose supports file merging. You keep a base file with the universal config and per-environment overrides:
# docker-compose.yml (base — committed)
services:
app:
image: ghcr.io/surya/myapp:${IMAGE_TAG:-latest}
environment:
APP_ENV: ${APP_ENV}
DATABASE_URL: ${DATABASE_URL}
LOG_LEVEL: ${LOG_LEVEL:-INFO}
ports: ["8000:8000"]# docker-compose.dev.yml (committed)
services:
app:
build: . # rebuild from local Dockerfile
environment:
APP_ENV: development
LOG_LEVEL: DEBUG
volumes:
- ./src:/app/src # live-reload
db:
image: postgres:16
environment:
POSTGRES_PASSWORD: dev
ports: ["5432:5432"]# docker-compose.prod.yml (committed — but ONLY references, no values)
services:
app:
image: ghcr.io/surya/myapp:${IMAGE_TAG}
environment:
APP_ENV: production
LOG_LEVEL: WARNING
deploy:
replicas: 3
resources:
limits: { cpus: "1", memory: "512M" }Invoke with the right combination:
docker compose -f docker-compose.yml -f docker-compose.dev.yml up docker compose -f docker-compose.yml -f docker-compose.prod.yml up -d
Later files override earlier. The actual values (DB passwords, API keys) come from the shell environment or a secrets manager — never from any of these files.
6. Branching Strategy → Deployment
Map git refs to environments so deploys are mechanical:
| Git ref | Deploys to | Approval |
|---|---|---|
| Feature branches | Nothing (CI tests only) | — |
| Pull request | Optional: ephemeral preview env (Render, Vercel, Fly) | — |
main | Staging — automatically on green CI | None |
Tag v* (e.g. v1.2.3) | Production — manual approval gate | Required |
The discipline:
- Never deploy from a branch other than
main— if it's not onmain, it hasn't been reviewed and CI hasn't fully tested it. - Production tags are immutable. Push a tag, deploy that exact SHA. Never re-tag.
- Promote, don't rebuild. The Docker image that deployed to staging is the exact same image that goes to prod — same digest, just retagged. No "let me kick off a fresh build for prod."
GitHub Actions for the deploy half:
deploy-staging:
if: github.event_name == 'push' && github.ref == 'refs/heads/main'
needs: [test, build]
environment: staging # GitHub "environment" with its own secrets
runs-on: ubuntu-latest
steps:
- run: ./scripts/deploy.sh staging ${{ github.sha }}
deploy-production:
if: startsWith(github.ref, 'refs/tags/v')
needs: [test, build]
environment:
name: production
url: https://app.example.com
runs-on: ubuntu-latest
steps:
- run: ./scripts/deploy.sh production ${{ github.ref_name }}The environment: block is GitHub's mechanism for deployment gating — repo settings → Environments → "production" → require manual approval before this job runs. Press the green button after eyeballing the diff; deploy proceeds.
7. Database Migrations Across Environments
Schema changes need a deployment story. The pattern:
1. In CI, on every push: run migrations against a throwaway DB to verify they apply cleanly.
2. In staging deploy: apply migrations automatically before the new app image takes traffic.
3. In production deploy: apply migrations as a separate, gated step. Often run by-hand or via a dedicated migration job, not auto-applied with the deploy.
Why the asymmetry? Schema changes can lock tables, break running queries, or fail mid-way. In staging, that's annoying. In prod, that's downtime. A separate step lets you:
- Review the SQL the migration will run (
alembic upgrade head --sql). - Schedule it for a low-traffic window.
- Roll back the app to the previous version first if the migration fails.
For zero-downtime: never write a migration that's incompatible with the current production code. The pattern is multi-step — add nullable column, deploy code that writes to it, backfill, deploy code that reads from it, drop the old column in a future deploy. Slow, deliberate, no outages.
8. Feature Flags
A feature flag is config that toggles a code path without a redeploy. The minimum:
class Settings(BaseSettings): feature_new_search: bool = False feature_realtime_updates: bool = False # in code if settings.feature_new_search: return new_search(query) else: return legacy_search(query)
For more than a handful of flags, a real flag service: LaunchDarkly, GrowthBook, Flagsmith, Unleash. These let you target flags per user, percentage-roll out, A/B test, all without redeploying.
The discipline: flags are temporary. A flag older than 6 months is a TODO everyone has forgotten. Schedule a quarterly flag-audit — remove the ones whose rollout is complete, ship the on-state permanently.
9. Logging Per Environment
The same code logs differently based on env:
| Env | Format | Level | Destination |
|---|---|---|---|
| Development | Human-readable, coloured | DEBUG | stdout |
| Staging | JSON structured | INFO | stdout → log aggregator |
| Production | JSON structured | INFO or WARNING | stdout → log aggregator (Datadog, Logflare, CloudWatch) |
Pattern with structlog:
import structlog, logging, sys, os def configure_logging(): level = os.environ.get("LOG_LEVEL", "INFO") logging.basicConfig(stream=sys.stdout, level=level, format="%(message)s") processors = [ structlog.stdlib.add_log_level, structlog.processors.TimeStamper(fmt="iso"), ] if os.environ.get("APP_ENV") == "development": processors.append(structlog.dev.ConsoleRenderer()) # coloured, human else: processors.append(structlog.processors.JSONRenderer()) # one JSON object per line structlog.configure(processors=processors)
setup added so this can run · defines
import os # noqa: F401 os.environ.setdefault("LOG_LEVEL", "example-log-level") os.environ.setdefault("APP_ENV", "example-app-env")
JSON-per-line is the universal aggregator-friendly format. Datadog, Loki, Splunk, CloudWatch — all parse it natively.
10. Error Tracking — Sentry, Per Environment
Production errors are different from dev errors. Sentry (or Rollbar, Honeybadger) gives you a per-environment view with stack traces, breadcrumbs, and user context:
import sentry_sdk sentry_sdk.init( dsn=settings.sentry_dsn, environment=settings.app_env, # tags every error with "staging" / "production" release=settings.git_sha, # tie errors to a specific deploy traces_sample_rate=0.1 if settings.app_env == "production" else 1.0, )
setup added so this can run · defines settings
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) settings = _AutoMock('settings')
The environment tag is the magic. In Sentry's UI you filter environment:production and dev errors disappear. Errors tagged with the git SHA tie back to a specific deploy — "this exception started after deploy v1.2.3" tells you what to roll back.
Don't ship sentry_dsn for dev unless you actually want dev errors in the same Sentry project. Usually you don't.
11. Health and Readiness Probes
Two endpoints, both required for serious deployment:
# /health — am I alive? @app.get("/health") def health(): return {"status": "ok"} # /ready — am I ready to serve traffic? @app.get("/ready") def ready(): try: db.execute("SELECT 1") cache.ping() except Exception as e: return {"status": "not_ready", "reason": str(e)}, 503 return {"status": "ready"}
setup added so this can run · defines app, db, cache
# Lightweight mock for objects whose attributes/methods aren't critical class _AutoMock: def __init__(self, name='mock'): self._name = name def __getattr__(self, k): return _AutoMock(self._name + '.' + k) def __call__(self, *a, **kw): print('-> ' + self._name + '() called') return _AutoMock(self._name + '()') def __repr__(self): return '<mock ' + self._name + '>' def __str__(self): return '<mock ' + self._name + '>' def __bool__(self): return True def __iter__(self): return iter([]) def __len__(self): return 0 def __getitem__(self, k): return _AutoMock(self._name + '[...]') def __setitem__(self, k, v): pass def __enter__(self): return self def __exit__(self, *a): return False async def __aenter__(self): return self async def __aexit__(self, *a): return False def __add__(self, o): return self def __radd__(self, o): return self def __sub__(self, o): return self def __mul__(self, o): return self def __rmul__(self, o): return self def __truediv__(self, o): return self def __eq__(self, o): return isinstance(o, _AutoMock) def __hash__(self): return hash(self._name) def __lt__(self, o): return True def __le__(self, o): return True def __gt__(self, o): return False def __ge__(self, o): return False def __mro_entries__(self, bases): return (object,) app = _AutoMock('app') db = _AutoMock('db') cache = _AutoMock('cache')
/health— minimal. Returns 200 if the process is alive. Used by Docker/Kubernetes to decide "restart this container?"/ready— checks downstream dependencies (DB, cache, external API). Used by load balancers to decide "route traffic here?"
The distinction matters: a container can be alive (process running) but not ready (DB connection still being established at startup). Both checks let the orchestrator do the right thing.
12. Deployment Strategies — Beyond "Push and Restart"
The simplest deploy: stop the old container, start the new one. Brief outage. Acceptable for hobby projects, never for paying customers.
Rolling deploy — replace containers one at a time. No outage, but new and old versions run simultaneously during the rollover. Your code must tolerate this (mostly: be backwards-compatible with the previous version's database schema and message formats).
Blue/green — two identical production environments. "Blue" is live; you deploy the new version to "green," smoke-test it, then flip the load balancer. Rollback is a second flip. Doubles your infra cost during the deploy window.
Canary — deploy to a small % of traffic (1%, then 5%, then 25%), monitor error rates and latency, expand if healthy. The most robust pattern; needs traffic-shifting infrastructure (Istio, Linkerd, ALB weighted target groups).
Pick by stakes — hobby app: rolling is plenty. Paying customers: blue/green or canary.
13. Observability per Environment
| Env | Metrics | Logs | Traces |
|---|---|---|---|
| Dev | None / Prometheus locally | print + structlog colour | Optional, mostly off |
| Staging | Prometheus + Grafana, retention 7 days | Aggregated (Loki, Datadog) | Sampled (10%) |
| Production | Prometheus + Grafana, retention 30-90 days | Aggregated, alerts wired | Sampled (1-5%) |
Staging gets observability so you can find issues there before prod. If staging has none, it can't pre-empt anything. Production gets it for incident response.
Alerts wire only in production. Staging alerts to a low-priority channel (or just a dashboard) — paging engineers because staging is down on a Sunday is how you burn out your team.
Common Mistakes
1. Sharing a database between staging and prod. The first developer to run a migration against the wrong env nukes real customer data. Always separate databases — different host, different name, different credentials.
2. No staging at all — "deploy and pray." Production becomes your test environment. Customers become your QA team. Fix: spin up a second deploy target today; even a $5/month Render service is staging.
3. Editing prod config by hand. Someone shells in, exports an env var, "fixes it." Six months later nobody remembers the fix, the box gets replaced, the fix evaporates. Config changes go through the same review/audit path as code: PR → merge → redeploy.
4. Hardcoded URLs for "the API." https://api.example.com in code means prod-hits-prod even when running staging. Either env-var-ise it or compute it from APP_ENV.
5. Feature-flag drift. Flags created for a 2-week rollout, still in code 18 months later. Audit quarterly; delete the flag once the rollout is done. Long-lived flags are bugs in waiting.
6. Different Docker images per environment. "Staging build" and "production build" with different code paths. The whole point of containers is the same artefact runs everywhere. Build once, promote across envs.
7. Auto-applying migrations in prod. A bad migration is now a deployment outage. Decouple: app deploy applies migrations in staging automatically, but in prod they're a separate, gated step.
8. Same Sentry project for all envs. Dev exceptions drown out prod incidents. Either separate projects or use the environment: tag to filter — and configure Sentry's alert rules to only fire on environment:production.
🎯 Your Turn — A Config Class Hierarchy Per Environment
Design a Config class hierarchy in plain Python (no Pydantic dependency) that:
1. Has a BaseConfig with universal defaults (PORT = 8000, LOG_LEVEL = "INFO").
2. Has three subclasses: DevConfig, StagingConfig, ProductionConfig.
3. Each subclass overrides only what differs (e.g. DevConfig.DEBUG = True, ProductionConfig.LOG_LEVEL = "WARNING").
4. A Config.from_env() classmethod reads APP_ENV and returns an instance of the correct subclass.
5. Required secrets (e.g. DATABASE_URL) are read from os.environ and the app refuses to start in staging/prod if they're missing — but in dev a local default is acceptable.
Skeleton:
import os from dataclasses import dataclass @dataclass class BaseConfig: # TODO 1: universal defaults ... @classmethod def from_env(cls): # TODO 4: dispatch on APP_ENV ... class DevConfig(BaseConfig): ... class StagingConfig(BaseConfig): ... class ProductionConfig(BaseConfig): ... config = BaseConfig.from_env()
Hint 1 — Dispatch table
A dict{"development": DevConfig, "staging": StagingConfig, "production": ProductionConfig} keyed by APP_ENV. Default to DevConfig when APP_ENV is unset (you're probably running locally).
Hint 2 — Required-secret guard
After instantiation, validate. InStagingConfig.__post_init__ / ProductionConfig.__post_init__, check that database_url isn't the placeholder and raise SystemExit with a clear message if it is. Don't crash dev — fall back to a local default there.
Show full solution
import os from dataclasses import dataclass, field @dataclass class BaseConfig: # Universal defaults app_env: str = "development" port: int = 8000 log_level: str = "INFO" debug: bool = False database_url: str = "" @classmethod def from_env(cls): """Dispatch on APP_ENV to the right subclass.""" env = os.environ.get("APP_ENV", "development").lower() subclass = _REGISTRY.get(env) if subclass is None: raise SystemExit( f"unknown APP_ENV={env!r}; expected one of {list(_REGISTRY)}" ) return subclass() def _require(self, name: str) -> str: """Read a required env var or fail fast.""" value = os.environ.get(name) if not value: raise SystemExit( f"[{self.app_env}] missing required env var: {name}" ) return value @dataclass class DevConfig(BaseConfig): app_env: str = "development" debug: bool = True log_level: str = "DEBUG" database_url: str = field( default_factory=lambda: os.environ.get( "DATABASE_URL", "postgres://localhost/app_dev" ) ) @dataclass class StagingConfig(BaseConfig): app_env: str = "staging" debug: bool = False log_level: str = "INFO" database_url: str = "" def __post_init__(self): self.database_url = self._require("DATABASE_URL") @dataclass class ProductionConfig(BaseConfig): app_env: str = "production" debug: bool = False log_level: str = "WARNING" database_url: str = "" def __post_init__(self): self.database_url = self._require("DATABASE_URL") if "localhost" in self.database_url: raise SystemExit( "production refused to start: DATABASE_URL points at localhost" ) _REGISTRY = { "development": DevConfig, "staging": StagingConfig, "production": ProductionConfig, } # --- Demo --- if __name__ == "__main__": # Simulate different environments os.environ["APP_ENV"] = "staging" os.environ["DATABASE_URL"] = "postgres://staging-db.internal/app" config = BaseConfig.from_env() print(type(config).__name__) # StagingConfig print(config.app_env) # staging print(config.log_level) # INFO print(config.debug) # False print(config.database_url) # postgres://staging-db.internal/app
setup added so this can run · defines
import os # noqa: F401 os.environ.setdefault("APP_ENV", "example-app-env") os.environ.setdefault("DATABASE_URL", "postgresql://user:password@localhost:5432/example")
What this hierarchy gives you:
- Defaults flow downhill — every subclass starts with
BaseConfig's values and overrides only what's different. ReadingProductionConfigis a 5-line file; no duplication of universal settings. - Environment-specific guardrails — production refuses to start if
DATABASE_URLpoints atlocalhost(a real bug that ships otherwise). Dev never has that guard; localhost is the expected default. - One-line dispatch —
BaseConfig.from_env()is the only call your app needs. Noif env == "production": ProductionConfig() elif ...scattered through code. - Fail-fast — missing required vars raise
SystemExitat startup with a clear message naming the env and the variable. NoKeyErrormid-request three hours into prod traffic.
For real applications, swap the dataclass internals for pydantic-settings — same idea, but with type coercion, validation, nested models, and automatic .env loading. The class hierarchy stays.
What You Learned
- Three environments minimum: dev (laptop), staging (mirrors prod), production. Staging must be safe to break.
- Anything that varies between environments goes in env vars. Code is identical across envs. Same Docker image, same git SHA.
- Env-var hierarchy: code defaults <
.env.example<.env< shell env < host secrets. Highest wins. docker compose -f base.yml -f prod.ymlmerges layered compose files — base + overrides per env.- Branching strategy maps to environments:
main→ staging (auto), tagv*→ production (gated). - Promote images, don't rebuild. The same image digest goes from staging to prod.
- Migrations: auto-apply in staging, gated in production. Multi-step backwards-compatible schemas for zero-downtime.
- Feature flags for safe rollouts; audit quarterly to delete dead flags.
/healthvs/ready— liveness vs readiness. Both required for serious deploys.- Sentry tagged with
environmentkeeps dev/staging/prod errors separated. Alerts wire only in prod. - Logging per env: human-readable in dev, JSON to an aggregator in staging/prod.
- Deploy strategies: rolling for most apps, blue/green or canary for high-stakes.
Next: Secrets Management — the credentials your environments share, where to store them, and how to rotate them when (not if) they leak.