Runbooks nobody follows are just documents
Why my first six runbooks failed, and the rewrite that cut incident resolution time roughly in half.
2 min readincident · runbooks
Writing
Case studies, how-tos and lessons learned. Mostly infrastructure, automation and applied AI — written down so the next person doesn't repeat the same four weekends.
Newest first
Why my first six runbooks failed, and the rewrite that cut incident resolution time roughly in half.
2 min readincident · runbooks
A phased migration that kept the site up every weekend, what the strangler-fig pattern actually looks like in practice, and the two things that nearly went wrong.
2 min readmigration · php
Six months of client work, tracking which tools stuck and which quietly stopped being opened. Includes the boring ones that won.
2 min readai · tooling
A short post about the incident that taught me to check the physical layer first, and the monitoring rule I added afterwards.
1 min readincident · dns
A restoration drill is the only backup that has ever saved anyone. How to run one in an afternoon, and the three failures it will find.
2 min readbackups · infrastructure
A concrete architecture for document extraction that a small Nepali services business can run on a budget, including the parts that go wrong.
2 min readai · llm
Latency numbers, D1 versus Postgres from Kathmandu, and the three assumptions that turned out to be wrong when I moved a client onto Cloudflare Workers.
2 min readcloudflare · workers
A finance team was losing four hours every Friday to a manual export. Here is exactly how we cut it to four minutes, and the part that mattered more than the automation.
2 min readautomation · python