Incident Postmortem Template: Blameless and Useful
A postmortem is written after the incident is over, to find out why it happened and what will stop it happening again. The six sections, a template to copy, a filled example, how to keep it blameless, and how action items become tracked work that actually gets done.
9 min read
An incident postmortem template has six sections: a summary a busy reader can stop after, the impact in numbers, a timeline with times, the root causes and contributing factors, what went well (and where you got lucky), and action items, each with one owner, a priority and a tracking ref. Write it blameless: it describes what the system allowed, not who slipped. Draft it within a few days while memories are fresh, review it with everyone who took part, and then treat the action items as ordinary work with a date to check on them. The template and a filled example are below.
A postmortem comes after the incident. The steps someone follows while it is happening belong in a runbook, and the special case of an AI agent causing the damage has its own checklist in AI agent incident response.
What a postmortem is for
Google’s SRE book, in its chapter on postmortem culture (opens in a new tab), defines a postmortem as a written record of an incident, its impact, the actions taken to mitigate or resolve it, the root causes, and the follow-up actions to prevent it recurring. The same chapter lists common triggers for writing one: user-visible downtime or degradation beyond a threshold, data loss of any kind, an on-call engineer having to intervene with a rollback or rerouting, a resolution time above a threshold, and a monitoring failure. It also says to agree the triggers before an incident, so nobody has to argue about whether this one counts.
The same document goes by other names. PagerDuty’s guide to postmortems (opens in a new tab) lists learning review, after-action review, incident review, incident report, post-incident review and root cause analysis as terms organizations use for it. A post-incident review template and a postmortem template are the same thing; use whichever word your team already says.
The blameless principle
The SRE book is specific about what blameless means: the postmortem must focus on identifying the contributing causes of the incident without indicting any individual or team, and it assumes everyone involved had good intentions and did the right thing with the information they had. The reason is practical. Where people are shamed for mistakes, they stop reporting them, and the next incident is found by customers instead. As the chapter puts it, you cannot fix people, but you can fix systems and processes.
- Write about actions and information, not people: “the deploy ran without the config check” rather than “Sam skipped the config check”.
- Ask why the action made sense at the time. If a reasonable person with the same dashboards would have done the same, the dashboards are the finding.
- Name roles in the timeline (“on-call engineer”, “support lead”) if names make people defensive. Keep names where they help coordinate follow-up.
- Be wary of any action item that reads “be more careful” or “remember to”. It asks a person to change and leaves the system as it was.
The six sections
- Summary: three or four sentences. What broke, for whom, for how long, the cause in one line, and the most important fix. Many readers will stop here.
- Impact: numbers. Users or accounts affected, requests failed, orders lost or delayed, revenue if you know it, support tickets, data lost or at risk. Say which figures are estimates.
- Timeline: timestamps with a time zone, from the first cause (often a change made hours or days before) through detection, response, mitigation and all-clear. Mark detection and mitigation clearly; the gaps between them are findings.
- Root causes and contributing factors: the trigger that set it off, the underlying cause that made the trigger dangerous, and the factors that made it worse or slower to fix. Most incidents have more than one, so resist the single-cause story.
- What went well, and where we got lucky: what to keep doing, and what saved you that you cannot count on next time. The lucky list is often where the best action items come from.
- Action items: each one a specific change, with a single owner, a priority, a tracking ref and a way to tell it is done.
An incident postmortem template you can copy
# Postmortem: [short description of the incident] Date of incident: [Month day, year] Duration: [start to end, time zone] Severity: [your scale] Status: Draft | In review | Final Postmortem owner: [one name] Written: [Month day, year] People involved: [names or roles] ## Summary [3-4 sentences: what broke, who was affected, for how long, the cause in one line, the most important fix.] ## Impact - Users / accounts affected: [number, and how you know] - Failed requests, orders, jobs: [number] - Data lost or at risk: [none / what] - Support tickets: [number] - Customer communication sent: [when, where] ## Timeline ([time zone]) [time] [the change or event that started it] [time] Detected: [by alert / customer / person] [time] [response step] [time] Mitigated: [what stopped the impact] [time] Resolved: [all clear, and how it was confirmed] ## Root causes and contributing factors Trigger: [what set it off] Root cause: [why the trigger could do this much damage] Contributing factors: - [what made it worse, slower to detect, or slower to fix] ## What went well - [keep doing this] ## Where we got lucky - [what saved us that we cannot count on next time] ## Action items | Action (verifiable) | Type | Owner | Priority | Ref | |---------------------|------|-------|----------|-----| | [specific change] | prevent / detect / mitigate | [one name] | [P1-P10] | [task ref] | ## Follow-up Review meeting: [date] Action items checked on: [date]
A filled example
An invented online store. The numbers are placeholders; the shape and the language are the point.
# Postmortem: Card payments failed at checkout for 47 minutes Date of incident: October 14, 2026 Duration: 10:02-10:49 ET Severity: SEV-1 Status: Final Postmortem owner: Priya (payments) Written: October 16, 2026 ## Summary For 47 minutes, every card payment failed at checkout. A config change renamed the environment variable holding the payment provider key; the app started without it and returned "payment unavailable". It was found by customer emails, not alerts. The key fix is a startup check that refuses to boot without the key. ## Impact - About 1,900 checkout attempts failed (from logs) - About 310 customers retried later and paid; the rest unknown - No data lost. No duplicate charges. - 64 support emails. Status page updated at 10:31. ## Timeline (ET) 09:40 Config change merged; review focused on another file 10:02 Deploy completes. Payment errors begin. 10:19 Detected: support lead forwards customer emails to on-call 10:27 On-call finds "missing PAYMENT_KEY" in app logs 10:44 Variable restored; app restarted 10:49 Resolved: test purchase succeeds; error rate back to normal ## Root causes and contributing factors Trigger: a variable was renamed in config but not in the app. Root cause: the app started and served traffic without a required secret instead of failing at startup. Contributing factors: - No alert on payment failure rate; detection took 17 minutes. - The deploy smoke test loads checkout but does not pay. ## What went well - Support escalated within minutes of the first emails. - The log message named the missing variable exactly. ## Where we got lucky - It happened on a Wednesday morning with the payments team online. ## Action items | Action | Type | Owner | Pri | Ref | |------------------------------------------|----------|--------|-----|---------| | App refuses to start without PAYMENT_KEY | prevent | Marcus | P1 | BUG-402 | | Alert: payment failures > 5% for 3 min | detect | Priya | P2 | ENH-403 | | Smoke test makes a test-mode payment | detect | Dana | P3 | ENH-404 | | Config renames need a check in CI | prevent | Marcus | P4 | ENH-405 | ## Follow-up Review meeting: October 16, 2026 Action items checked on: October 30, 2026
Notice what is absent: the name of the person who merged the config change. The finding is that a reviewer could approve it and the app could start without the key. Both are true whoever merged it, and both are fixed by the action items.
Action items that get done
The SRE workbook’s chapter on postmortem culture (opens in a new tab) walks through a bad postmortem and a good one, and its criticism of the bad one reads like a checklist. Its action items were mostly mitigating rather than preventing, all had the same priority so nobody knew what to do first, two used vague verbs like “improve”, and only one had a tracking bug. Without a formal tracking process, it says, action items from postmortems are often forgotten, resulting in outages. The good example gives every action item an owner and a tracking number, a priority, and a verifiable end state.
- One owner per item. A single name who will see it through, with others helping. The workbook says the same of the postmortem itself.
- Verifiable. “Alert when payment failures exceed 5% for 3 minutes” can be checked; “improve monitoring” cannot.
- Prevent, detect or mitigate. Label each; a list with no prevent items means the same incident is expected again.
- Priority relative to other work, not a flat “high”. If everything is P1, the list gives no order.
- Tracked where the team already works. An action item that lives only in the postmortem document is not on anyone’s list.
Reviewing and following up
PagerDuty’s published internal policy is to complete postmortems within 3 calendar days for a Sev-1 and 5 business days for a Sev-2, and to prioritize the postmortem over planned work. Pick deadlines that suit your team, but pick them. Then hold one short review meeting with everyone who took part, walk the timeline, agree the causes and the action items, and mark the document final.
The follow-up is where most postmortems fail. Put a date two weeks out to check the action items, and read the list aloud: done, in progress, or not started and why. The SRE workbook warns against rewarding the writing of postmortems but not the closing of their action items, which produces a pile of documents and no fewer outages. For security incidents, NIST’s SP 800-61 Revision 3 (opens in a new tab), published in April 2025, frames incident response within the Cybersecurity Framework 2.0; a blameless postmortem with tracked action items fits inside it rather than replacing it.
Turning action items into fenbs tasks
fenbs does not hold the postmortem document; keep that in your docs. What it holds is the work that comes out of it. Each action item becomes one task: a bug with a BUG- ref when something was broken, like the startup check above, and an enhancement with an ENH- ref when something was missing, like the alert. Priority runs 1 to 10, 1 the most urgent, so the prevent item that stops a repeat can sit above routine work. fenbs has no assignee field you can set, so write the owner as the first line of the note, with the postmortem’s title and link, and the task’s History shows who moved it and when.
The quickest route is Add many, next to + New task: paste one line per action item, starting each with bug: or enhancement:, and indented lines become that task’s note. When an item is done, its test status records how the fix was checked, Tested, Partly tested or Failed, with notes. Filtering the board for tasks that are Completed but not tested is the follow-up question in one view. And when the postmortem produces a standing instruction, such as “the app must refuse to start without a required secret”, record it as a rule on the Decisions and rules page, with the person who decided it named as the decider; every connected AI assistant reads the rules first.
- bug: App starts without PAYMENT_KEY instead of refusing to boot
Owner: Marcus. From postmortem "Card payments failed at
checkout", October 14, 2026. Done = boot fails with a clear error.
- enhancement: Alert when payment failures exceed 5% for 3 minutes
Owner: Priya. Same postmortem.
- enhancement: Deploy smoke test makes a test-mode payment
Owner: Dana. Same postmortem.Related
The steps to follow during the incident: runbook template. When an AI agent caused it: AI agent incident response. Writing each fix up properly: bug report template. Recording the rules that come out of it: decision log template.