The 3AM Runbook
What actually happens when the pager goes off, and how to make future-you calmer than present-you.
- #on-call
- #incident-response
- #observability
// updated 2026.05.09
The pager doesn’t wake you because something broke. It wakes you because something broke and nobody knew how to fix it fast. The first part you can’t always prevent. The second part you can.
The first ninety seconds
Before you touch anything, answer three questions.
Is it us or is it upstream? Check the dependency dashboard first. A good chunk of “our” outages turn out to be someone else’s region having a bad day.
What changed? Last deploy, last config push, last cert rotation. When you’re half awake, the timeline of recent changes is worth more than any log. Most incidents are something you did, not something that decided to break on its own.
Can I stop the bleeding without understanding it yet? Roll back, shed load, flip a flag. Get things stable, then figure out why. Nobody remembers that you took nine minutes to find the root cause. They remember that the site was down for nine minutes.
The runbook writes itself if you let it
Every incident should end with one question. What would have made this five minutes shorter? A dashboard link, a rollback command, an alert that fires earlier. Whatever it is, it goes in the runbook before the retro is over, while you still remember how it felt.
# the command I wished I'd had at 3:07am
kubectl rollout undo deployment/api --to-revision=$(
kubectl rollout history deployment/api | awk 'END{print $1-1}'
)
Good runbooks are boring. They’re a checklist a tired person can follow without thinking. If yours needs you to be clever, it’ll fail at exactly the moment you have the least cleverness to spare, which is the whole reason you’re reading it at 3am in the first place.
Back to Case Files