Every team has a runbook. Most were written on a quiet Friday afternoon by someone who no longer works there. The steps reference a dashboard that moved, a config flag that was renamed, and a contact list with two people who left last quarter.
Nobody notices until the pager goes off at 2 AM.
The on-call engineer opens the doc, skims it, realizes step three no longer applies, improvises steps four through seven, and resolves the incident through tribal knowledge and adrenaline. In the morning, someone adds a comment: "Update this." Nobody does.
This is the lifecycle of most runbooks. Written once, bookmarked, left to rot.
Runbooks don't fail because people are lazy. They fail because they exist outside the flow of work.
Think about what gets maintained in your codebase. Tests run on every commit. Linters enforce style on every pull request. If something breaks, the build breaks, and someone fixes it before merging. Maintenance is automatic because it's embedded in a process that already happens.
Now think about your runbook. It lives in a wiki. It has no connection to your deploy pipeline, your alerting rules, or your release cycle. Nothing forces anyone to open it unless something is already on fire. And when something is on fire, you don't have time to fix the documentation — you have time to fix the fire.
The problem is structural, not cultural. Runbooks decay because nothing prevents them from decaying.
A runbook that nobody validates is like a fire extinguisher that nobody inspects. It looks reassuring on the wall. It might work. But the only way to know is to check it on a schedule — not to wait for the fire.
Building inspectors don't trust that extinguishers work because they were functional at install. They check them periodically, tag them with a date, and replace them when they expire. The inspection is cheap. A dead extinguisher during a real fire is not.
Your runbook needs an inspection cycle. Not a big one. A small, boring, repeatable one.
The practice: attach runbook validation to your deploy checklist.
Every team already has some kind of pre-release checklist. It might be formal or it might be three bullet points in a shared channel. Either way, add one item: "Can the on-call engineer follow the runbook for this service without asking anyone a question?"
That's the test. Not "is the runbook complete?" Not "is it well-written?" Just: can someone who wasn't in the room when the change was made actually follow the steps?
This works for three reasons.
First, it happens at the right time. A deploy is the moment most likely to introduce a gap between what the runbook says and what the system does. Checking at deploy time catches drift before it matters.
Second, it's small. You're not asking anyone to rewrite documentation. You're asking them to read a page and confirm it still matches reality. If it doesn't, fix the two lines that changed. Five minutes.
Third, it assigns responsibility to the person who knows what changed. The engineer shipping the deploy knows which flags moved, which endpoints shifted, which dependencies were added. Asking them to glance at the runbook is the smallest possible version of knowledge transfer.
A runbook that gets validated every release cycle starts to look different from one written once.
It gets shorter. Steps that nobody follows get removed instead of lingering. Assumptions get stated explicitly because someone had to verify them last week. Contact lists stay current because someone noticed the wrong name during the last check.
Honest runbooks have dates on them — not "last edited" timestamps buried in wiki metadata, but a visible line that says when the doc was last confirmed accurate and by whom. That line is the tag on the fire extinguisher.
They also separate what changes often from what stays stable. The parts that drift every few releases get pulled to the top where they're easy to scan. The parts that haven't changed in a year settle to the bottom.
The instinct when runbooks fail is to launch a documentation project. Designate an owner. Schedule a review meeting. Build a template. Create a tracker.
Most of these projects produce a burst of updated docs and then die within two quarters. The meetings stop. The tracker goes stale. You're back where you started, just with newer timestamps.
Operational resilience doesn't come from projects. It comes from habits — small actions repeated at a cadence that already exists. Tying runbook checks to deploys works because deploys already happen. You're not adding a process. You're adding a line item to a process you already follow.
The goal is not perfect documentation. The goal is documentation that doesn't lie to you at 2 AM.
Be the first to comment.
0 comments
Loading comments...