Runbooks nobody follows are just documents
Why my first six runbooks failed, and the rewrite that cut incident resolution time roughly in half.
2 min read
I wrote my first runbook at a company where I was the only person with permissions to act. It was thorough, correct, and never opened, because there was one of me and it was not me reading it.
Why the first six failed
- 01
They were written after the incident, from memory
Which meant they described what I remembered doing, not what I actually did. The gaps were always in the awkward middle.
- 02
They assumed the reader knew the system
"Check the queue" is useless to someone who has never seen the admin panel.
- 03
They had no stopping point
No indication of what a good outcome looks like, so the reader kept going long after the problem was solved.
- 04
They were not tested by anyone else
I had verified the commands worked. I had not verified a cold reader could follow them.
What the rewrite looks like
Every runbook now follows the same four-part structure. It is a template, not a suggestion.
1. Symptom
One line, in the words an alert would use. If I cannot match the runbook to the alert, the runbook is not linked from the alert and nobody will find it.
2. Confirm, before acting
The check that proves the runbook applies. Explicitly including the cheap check that could disprove it.
# Confirm before touching anything.
# 1. Is it actually this failure? (disprove-first, not confirm-first)
curl -fsS https://status.internal/health/live && echo "app alive"
# 2. Is it the queue specifically?
systemctl is-active worker-queue && echo "queue running"
# 3. How deep is the backlog, and since when?
workerctl stats --json | jq '{depth: .depth, oldest_age_s: .oldest_age}'
3. Act
Numbered steps. Each one a single action. Each one with an expected result, so a failed step is obvious rather than inferred.
# Step 1 — drain. Expect depth to fall to 0 within 60s.
workerctl drain --wait 60 --timeout 30
# Step 2 — verify. Expect {"depth":0,"oldest_age_s":<60}
workerctl stats --json | jq '{depth, oldest_age_s}'
# Step 3 — only if step 2 failed: restart. Expect active within 10s.
sudo systemctl restart worker-queue && sleep 10 && systemctl is-active worker-queue
4. Stop and escalate
The part I left out for years. Explicitly: what a good outcome is, what to do if it is not achieved, and who to wake.
The test that actually catches problems
Once a quarter, a colleague who did not write the runbook follows it against a staging clone while the author watches and stays silent.
That is uncomfortable. It is also the only review that finds the steps which only make sense if you already knew the answer.
What this changed about how I work
I write runbooks during the incident now, not after. It is faster, it is more accurate, and it produces a document that describes reality rather than recollection. I also accept that a runbook is unfinished until someone else has followed it.