Every team says they do post-mortems. Most teams do them once, produce a long document, and never look at it again. Or they skip the document entirely and hash it out in a chat thread that scrolls away by Friday.
Both approaches share the same failure: nothing changes after the incident. The point of reviewing an incident is not to produce a record. It is to produce a change. If your process does not reliably produce changes, your process is decoration.
We got tired of decoration. So we replaced our open-ended post-mortem format with five questions. The same five, every time, no exceptions.
A 30-page post-mortem feels thorough. It is also a trap. Long documents invite long debates about phrasing. They defer action behind "we need to finish the write-up." And once finished, they sit in a folder nobody opens because reading 30 pages is a commitment no one scheduled.
Short formats work better for the same reason checklists work in operating rooms: they lower the activation energy to actually do the thing. When the template fits on one screen, people fill it out the same day. When it demands specific owners and dates, people follow through or get visibly called out.
Rigid does not mean shallow. It means disciplined. You can go deep on each question. You just cannot wander.
Write a timeline a new hire could follow. No jargon, no hedging. Start with the first observable symptom and end with the moment service returned to normal. If you cannot explain what happened in a short paragraph, you do not yet understand what happened — and that is worth knowing before you move on.
The goal is shared reality. Everyone involved should read this summary and nod. If someone disagrees with the timeline, surface that disagreement now, not three weeks later.
This is the question most post-mortems skip or bury. It is the most important one. You are not asking "what broke?" You are asking "what signal existed before the break that we did not see or did not act on?"
Maybe an alert was too noisy and got ignored. Maybe a manual process had no verification step. Maybe someone noticed something odd on Monday and the incident happened on Wednesday. The gap between the available signal and the response is where your real lessons live.
Resist the urge to write "human error" here. Human error is a description, not a cause. If a person made a mistake, ask what about the system made that mistake easy to make and hard to catch.
Not "what could we change" or "what should we consider." What are we changing. Present tense, concrete.
This question forces a commitment. If the answer is "nothing," write "nothing" and explain why. That is legitimate — some incidents are true outliers that do not justify a process change. But writing "nothing" in plain text makes you own that decision instead of letting inaction hide behind ambiguity.
Good changes are small and specific. "Improve monitoring" is not a change. "Add an alert that fires when error rate exceeds X for Y minutes" is a change.
A change without an owner is a wish. Every item from question three gets a name next to it. One name, not a team. Teams do not ship fixes; people on teams ship fixes.
Social pressure does useful work here. When your name is next to a line item in a shared document, you are more likely to finish it. Not because of fear, but because of clarity. You know it is yours.
Pick a date. Put it on a calendar. On that date, someone — often the person who ran the review — looks at each change from question three and marks it done or not done.
This is the question that converts a review into a habit. Without a re-check, you are relying on memory and good intentions. Memory fades. Good intentions get crowded out by next week's priorities.
We set the re-check for two weeks out. Long enough to ship a fix, short enough that the incident still feels real.
No severity ratings. No root-cause categories from a dropdown menu. No sign-off chain. Those artifacts serve compliance, not learning. If you need them for regulatory reasons, add them as an appendix. Do not let them take over the body of the review.
We also skip blame assignments. The five questions are oriented around the system, not the individual. Question two asks "what did we miss" — we, collectively. Question four assigns ownership of the fix going forward, not responsibility for the failure.
You can tell whether your incident review process works by asking one question six months from now: did the same class of incident happen again?
If it did, your process produced documents. If it did not, your process produced change. Five questions, filled out honestly, with names and dates attached, will get you to the second outcome more often than any template with a table of contents.
Be the first to comment.
0 comments
Loading comments...