What happened
On 31 January 2017, GitLab.com's database replication had stopped, and an engineer was rebuilding the secondary server. That job needs an empty data directory on the secondary, so they deleted it. GitLab's post-mortem describes what went wrong next: they did so "errantly thinking they were doing so on the secondary. Unfortunately this process was executed on the primary instead."
They stopped the command "a second or two after noticing their mistake, but at this point around 300 GB of data had already been removed."
GitLab.com was down for about 18 hours, most of it spent copying data back over slow disks. Changes made in a six-hour window were gone for good: "roughly 5,000 projects, 5,000 comments and 700 new user accounts." Git repositories and wikis were not affected. Only the database was.
Why the backups didn't save them
Running the right command on the wrong server can happen to anyone. What turned it into lost data was that GitLab had several ways to recover, and when they needed them, almost none worked:
- The daily
pg_dumpbackups weren't there. "The S3 bucket was empty." The backup job used PostgreSQL 9.2 tools against a 9.6 database, so it failed with an error every time. The warnings were sent by email, and the emails were "rejected by the receiver". Nobody saw them. - Cloud disk snapshots covered other servers but not the databases, "as we assumed that our other backup procedures were sufficient."
- Replication was the thing being repaired, so the secondary had no copy to offer.
- What saved them was luck. A disk snapshot is normally taken every 24 hours to refresh a staging copy. That day an engineer had taken one by hand "roughly 6 hours before the outage", because they wanted fresher data for load testing. GitLab restored from that snapshot, which is why six hours of data were lost rather than everything.
The failure you can reproduce
The dangerous part wasn't the error. It was an error that nobody saw. Here's a small version: a backup step that fails, a wrapper that swallows the failure, and a job that reports success anyway:
import json def dump(db, tool_version, server_version): if tool_version != server_version: raise RuntimeError(f"server version mismatch: server {server_version}, tool {tool_version}") return json.dumps(db) def nightly_backup(db): try: return dump(db, tool_version="9.2", server_version="9.6") except Exception as e: # The warning goes somewhere nobody reads. Here: nowhere. return "" db = {"projects": ["api", "docs", "site"], "users": ["ada", "grace"]} backup = nightly_backup(db) print(f"backup job finished, {len(backup)} bytes written")
backup job finished, 0 bytes written
The job ran and finished, and nothing told anyone it had written nothing.
Test the restore, not the backup
The fix is to stop trusting that a backup exists and prove you can restore it: load it back, and compare what comes back with what you expected.
import json def restore_check(backup, expected): if not backup: raise RuntimeError("backup is empty") restored = json.loads(backup) for table, rows in expected.items(): got = len(restored.get(table, [])) if got != len(rows): raise RuntimeError(f"{table}: restored {got} rows, expected {len(rows)}") return "restore OK" db = {"projects": ["api", "docs", "site"], "users": ["ada", "grace"]} good = json.dumps(db) print(restore_check(good, db)) for bad in ["", json.dumps({"projects": ["api"], "users": []})]: try: restore_check(bad, db) except RuntimeError as e: print("ALERT:", e)
restore OK ALERT: backup is empty ALERT: projects: restored 1 rows, expected 3
A check like this fails loudly on the first night the backup breaks, not on the day you need it. GitLab's own action list includes the same idea at full scale: "Automated testing of recovering PostgreSQL database backups".
Fix the system, not the person
GitLab's post-mortem doesn't name the engineer, and its fixes aim at the situation they were in. In GitLab's words, the focus was "to improve disaster recovery, and making it more obvious as to what host you're using; instead of preventing production engineers from running certain commands."
Their list of fifteen follow-ups includes:
- changing the shell prompt on every server "to more clearly differentiate between hosts and environments", so primary and secondary no longer look alike
- monitoring the backups themselves
- hourly disk snapshots of the production databases, and cloud snapshots switched on
- "Assign an owner for data durability", so that "are the backups working?" is someone's job rather than everyone's assumption
What you can take from it
- Never swallow an error in a job that runs unattended. If the only place a failure can go is an inbox, check that the inbox receives it.
- Measure the output, not the exit. "Finished" and "0 bytes" can both be true.
- Restore on a schedule. A restore you've actually run is the only proof a backup works.
- Make the dangerous place look dangerous. Different prompts, colours or confirmation steps for production are cheap, and they work on tired people.
- Count your safety nets honestly. GitLab had four ways to recover. On the day, one worked, and only because of a snapshot someone had taken by chance.