Backups: the checks that matter are the ones nobody runs

A restoration drill is the only backup that has ever saved anyone. How to run one in an afternoon, and the three failures it will find.

2 min read

I have never once been called to a “restore the backup” incident where the backup was the problem. In every case the backup existed, was current, and had never been used.

The three failures you will find

1. The backup includes what you thought, but not what you need

The nastiest failure is a backup that contains the database, the config, and the application — and none of the uploaded files. Because uploads were on a different filesystem, on a different mount, added by a different developer, six months after the backup policy was written.

Write down what “restore” means before you write what to back up. Recovery objective first: how much data can you afford to lose, and how long can you be down?

2. It restores, but slowly enough to matter

A 40 GB restore over a consumer uplink took eleven hours. The team had a four-hour recovery target. The backup was technically fine and completely useless.

Measure restore throughput, not backup success. pv on a decompressed stream gives you the number that actually matters.

3. It restores, but not to a working state

This is the one that finds me the most. The database comes back. The application then cannot start, because the secrets were never part of the snapshot, or the migration that ran last month was not in the deployment artefact, or the file permissions assumed a user that no longer exists.

An afternoon drill

  1. 01

    Pick the smallest restore point that counts

    Usually "the last week". Restoring a year-old snapshot proves you can read a file.

  2. 02

    Restore to something isolated

    A spare VM or a container on the existing hypervisor. Never over production, never "just to check".

  3. 03

    Restore config and secrets separately

    Secrets come from a password manager, not the backup. If they came from the backup, that is a finding.

  4. 04

    Start the app and run a real check

    Log in, read a record, submit a form. Not a health endpoint — a health endpoint lies.

  5. 05

    Time every step, write it down

    The timeline is the deliverable. The restore is just the method of producing it.

Script the boring part so you actually run it

Manual drills do not happen. This is a skeleton; adapt the checks to your stack.

#!/usr/bin/env bash
set -euo pipefail

TARGET_HOST="${1:?usage: drill.sh <restore-host>}"
START=$(date +%s)

echo "→ fetching artefacts"
scp ./backup-$(date -u +%F).tar.gz "${TARGET_HOST}:/var/tmp/"
ssh "${TARGET_HOST}" 'cd /var/tmp && tar xzf backup-*.tar.gz'

echo "→ restoring database"
ssh "${TARGET_HOST}" 'sudo systemctl stop app && sudo -u postgres psql -f /var/tmp/db.sql app'

echo "→ fetching secrets from vault"
ssh "${TARGET_HOST}" 'sudo vault read -field=env secret/app/prod > /etc/app/env'

echo "→ starting app"
ssh "${TARGET_HOST}" 'sudo systemctl start app && sleep 8 && systemctl is-active --quiet app'

echo "→ functional check"
ssh "${TARGET_HOST}" 'curl -fsS https://localhost/health/real >/dev/null'

END=$(date +%s)
echo "✓ drill complete in $(( END - START ))s — log it"

What to log every time

Total wall-clock time, per-step timings, anything you had to fix manually, and every assumption that turned out to be wrong. After three drills the manual fixes become the automation backlog, and the timeline becomes the recovery objective you can actually promise a client.

On this page

Deploying a Cloudflare Worker for a Nepali SaaS, and what surprised me

Latency numbers, D1 versus Postgres from Kathmandu, and the three assumptions that turned out to be wrong when I moved a client onto Cloudflare Workers.

2 min readcloudflare · workers

Upgrading a decade-old PHP app without stopping the business

A phased migration that kept the site up every weekend, what the strangler-fig pattern actually looks like in practice, and the two things that nearly went wrong.

2 min readmigration · php