All posts
Runbooks the On-Call Engineer Will Actually Open at 3am
runbooksincident responseon-call managementSRE

Runbooks the On-Call Engineer Will Actually Open at 3am

Most runbooks are written for someone calm at a desk, not someone half-awake under pressure. What makes an on-call runbook get opened and used.

Yathartha Shekhar

Yathartha Shekhar

Founder, Fluidify.ai

October 9, 2026

7 min read

Key takeaways

  • Most runbooks are written by someone calm, at a desk, with full context, for a reader who is none of those things during an actual incident. That mismatch, not a lack of documentation effort, is why so many runbooks exist and don't get opened.
  • A runbook that gets used at 3am is scannable in seconds, structured as decisions and commands rather than prose explanation, and current enough that the on-call engineer trusts it over their own memory.
  • Staleness is the silent killer: a runbook that's wrong once teaches the team to stop trusting all of them, which is worse than having no runbook at all.

Why on-call runbooks don't get opened

A runbook is usually written right after an incident, or during a calm sprint dedicated to documentation, by someone who understands the system well and has time to explain it properly. It gets read, if it gets read at all, by someone at 3am, cognitively impaired by being woken up, under time pressure, trying to stop something from getting worse. Those are two completely different reading conditions, and a document written for the first one is very often unusable in the second. Not because it's wrong, but because it's shaped for the wrong reader.

The tell is a runbook that starts with background: why the system works the way it does, historical context, the reasoning behind a design decision. All of that is genuinely valuable at some point, just not at 3am, when the reader needs the fourth paragraph's worth of information in the time it takes to skim a heading.

None of this means runbooks aren't worth writing. Google's SRE workbook describes the norm in its on-call chapter: whenever an alert is created, a matching playbook entry usually is too, because good ones reduce stress, time to repair, and the risk of human error. The problem is the shape of the document, not the habit of writing one.

What makes an on-call runbook actually get opened

Scannable in seconds, not read start to finish: headings that are actions or symptoms ("Service returning 503s", not "Background on the service architecture"), so the on-call engineer can jump to the relevant section without reading the whole document to find it. A runbook that requires reading top to bottom to find the relevant part won't get read that way under pressure. It'll get abandoned in favor of guessing.

Commands, not descriptions of commands: "Run kubectl rollout restart deployment/api -n prod" beats "restart the API deployment in the production namespace." A tired person copy-pastes a command far more reliably than they correctly translate a description into one, and a typo in a hand-typed command during an incident is its own risk.

Decision structure, not narrative prose: "If X, do Y. If Y doesn't work, do Z." reads faster under pressure than a paragraph explaining the same logic. The runbook's job in the moment is to be a lookup table, not an essay.

One clear owner and one clear scope per runbook: a runbook trying to cover an entire service's worth of possible failures becomes a wall of text nobody can navigate quickly. Narrow runbooks, one per failure mode, linked from wherever that failure mode's alert fires, work better than a single sprawling document covering everything. If an alert fires often enough to need a runbook but nobody can say what action it calls for, that's an alert problem, not a documentation problem; see why alert fatigue survives every tooling migration.

The runbook staleness problem

A runbook that's wrong is worse than no runbook. The first time someone follows stale instructions and it makes things worse, or just wastes ten critical minutes, they stop trusting runbooks generally, not just that one. That erosion of trust is expensive to rebuild and cheap to avoid, mostly by keeping runbooks narrow enough that updating them isn't a big lift, and by treating a runbook that turned out to be wrong during a real incident as something the postmortem should explicitly fix, not something that quietly stays wrong until the next person hits it. The same SRE workbook chapter puts it plainly: playbook details go out of date as fast as the production environment changes.

A useful practice: whenever a runbook gets followed during a real incident, note in the postmortem whether it was accurate and current. If it wasn't, fixing it is a concrete action item, the same way fixing a bad alert rule would be. Keeping runbooks in the same repository as the alert rules, reviewed in the same pull request, helps for the same reason it helps with escalation policies as code: a change to one prompts a look at the other.

Where runbooks should actually live

A runbook that's hard to find in the moment is functionally the same as a runbook that doesn't exist. Linked directly from the alert that would trigger needing it, or from the incident record once one's declared, beats a wiki search that assumes the on-call engineer remembers the right search term while half-awake. The goal is zero navigation between "I got paged" and "I'm looking at the relevant runbook." Every extra click is a chance to give up and start improvising instead.

If you run Prometheus, this is built in: alerting rules have an annotations field meant for exactly this kind of longer information, and the Prometheus docs name runbook links as a use for it. The convention is a runbook_url annotation on each alert, which Alertmanager passes along with the alert to whatever pages you. The kube-prometheus runbooks are a good public example: every alert in that project points at its own short page, like this one for TargetDown.

A quick runbook audit

Pick a runbook and time how long it takes a fresh pair of eyes to find the first actionable command in it, starting from the top. If it's more than a few seconds, it's probably too front-loaded with context that belongs later in the document, or in a separate reference doc entirely, not in the thing someone's trying to use under pressure.

An on-call runbook template

A shape that keeps the actionable part on top:

  • Symptom: the alert name and what the person paging will actually see.
  • Impact: who or what is affected, in one line, so the reader can judge urgency.
  • First command: the one thing to run first, copy-pasteable.
  • If this, then that: two to five decision steps, each with its command.
  • Escalate when: the point at which to stop and page someone else, and who.
  • Background: architecture notes and history, at the bottom, for later.
  • Last verified: the date and incident when someone last followed it and it worked.

FAQ

What should an on-call runbook include? The symptom, the impact, the first command to run, a short decision tree, when to escalate, and the date it was last verified. Background belongs at the bottom, not the top.

Should runbooks be written by the person who built the system? They're the best starting source, but the person testing whether it's usable should ideally be someone less familiar with the system. Familiarity is exactly what makes a writer blind to steps that feel obvious to them and aren't to anyone else.

How often should runbooks be reviewed? Whenever the underlying system changes meaningfully, and whenever one gets used during a real incident, which is the best signal for whether it's actually still accurate. A calendar-based review cadence catches less than usage-triggered review does.

Is a runbook different from a postmortem? Yes. A postmortem explains what happened after the fact; a runbook tells someone what to do while something is still happening. Conflating the two formats, by adding runbook-style fix steps to a postmortem or narrative explanation to a runbook, weakens both.

Is a runbook different from a handoff note? Yes. A runbook covers a failure mode in general; a handoff note covers one specific incident's state at a shift change. Both matter at 3am. For the handoff side, see who owns the pager when the team is spread across time zones.

One disclosure: we build FluidifyAI Regen, and this is part of why it keeps every alert, annotations like runbook_url included, attached to the incident that alert opened, instead of assuming the on-call engineer remembers where to look.