I got paged for a DNS outage that was a laptop on hotel wifi

A short post about the incident that taught me to check the physical layer first, and the monitoring rule I added afterwards.

1 min read

Total outage. Every user. Roughly twenty minutes. The root cause was a laptop that had decided to run a local DNS resolver while connected to hotel wifi, and had briefly been reachable on a network segment it should not have been.

What I did, in the wrong order

I logged in and started auditing DNS records. Zone file, TTLs, serial numbers, the works. All correct. All unchanged. Twenty minutes gone.

The question I skipped was the cheapest one: who is answering?

# What is actually resolving this name right now?
dig +short example.com

# Who is answering, and is it the server I think it is?
dig +trace example.com | tail -20

# Compare against the public view, bypassing any local override
dig @1.1.1.1 +short example.com

dig @1.1.1.1 was the command that mattered. Local resolution returned a different answer than the public resolver, which told me the problem was below the authoritative servers and above my infrastructure.

What I changed afterwards

A resolution check that compares, not just resolves

Monitoring that asks “is this name resolving?” will tell you it resolved. It will not tell you it resolved to something wrong. The comparison is the point.

# Alert when internal and public resolution disagree.
LOCAL=$(dig +short +time=2 +tries=1 "$HOST" | head -1)
PUBLIC=$(dig @1.1.1.1 +short +time=2 +tries=1 "$HOST" | head -1)
[ "$LOCAL" != "$PUBLIC" ] && echo "split horizon for $HOST: $LOCAL vs $PUBLIC" | mail -s dns-split alerts@…

An escalation note that names the ladder

The runbook now says, in order: check the local resolver, check the physical layer, then touch configuration. I put it first because that is exactly where I did not start.

The honest part

Nothing about this was technically hard. It was a twenty-minute outage caused by a twenty-second check I did not run, during an incident where I felt pressure to do something visible. I opened a config file because opening a config file felt like progress.

The discipline is resisting that. On an outage, the first action should be the cheapest one that could disprove your leading hypothesis — not the most familiar one.

On this page

Runbooks nobody follows are just documents

Why my first six runbooks failed, and the rewrite that cut incident resolution time roughly in half.

2 min readincident · runbooks

The transaction succeeded and my internet got cut off

A renewal payment cleared at the bank, cleared at NCHL, and still could not be activated. What that episode taught me about digital payments reconciliation in Nepal.

2 min readdigital-payments · nepal