Writing Incident Reports & Postmortems
The work of incident response isn't finished when the threat is contained — it's finished when it's written up so the organisation learns from it. This is the most undervalued skill in IR and one interviewers probe directly ("tell me about a postmortem you wrote"). This file covers the documents you produce during and after an incident: how to structure them, the blameless principle, the language to use, and the templates to reuse.
Last verified2026-06
Why Writing Is a Core IR Skill
A brilliant investigation that's poorly communicated has a fraction of its value. Reports and postmortems are how:
on spend, on risk acceptance, on whether to notify customers/regulators.
so the same root cause doesn't recur.
engineering fixes the bug, IT closes the gap, detection builds the rule.
regulatory clocks (GDPR 72h) depend on clear, timely documentation.
clear, honest writing is how a security team earns credibility with the rest of the company.
Memory hook"if it isn't written down, it didn't happen." In IR, the investigation lives in the analyst's head until it's documented. The report is the deliverable. An incident with no written record can't drive change, can't be defended to a regulator, and will recur.
The Documents You Produce (and when)
| Document | When | Audience | Purpose |
|---|---|---|---|
| Incident log / timeline | During — live | The response team | The running record of what happened and what you did, with timestamps |
| Situation report (SITREP) | During — periodic | Leadership / stakeholders | Short status update: where we are, what we need (use SBAR) |
| Incident report | After containment | Management, sometimes customers/regulators | The formal "what happened, impact, what we did" record |
| Postmortem / RCA | After — the retro | Engineering + security, internal | Why it happened and how to prevent recurrence — the learning document |
| Customer / regulator notification | As required | External | Legally/contractually mandated disclosure |
The incident log feeds everything else — keep it meticulously during the event, because reconstructing a timeline from memory afterward is error-prone and slow.
The Blameless Principle (the single most important idea)
A blameless postmortem assumes that everyone acted reasonably given the information they had at the time, and focuses on systems and processes, not people. This isn't about being nice — it's about getting the truth.
Why blameless worksif engineers fear punishment, they hide details, and you never learn the real root cause. Psychological safety is what makes people say "I ran that command because the runbook was ambiguous" instead of staying silent. Blame drives the truth underground; blamelessness surfaces it.
Blameful (wrong): "Alex deployed bad config and took down auth."
Blameless (right): "A config change passed review and deployed to production
because the validation pipeline didn't check for this class
of error. Two reviewers also missed it, suggesting the
review checklist doesn't cover it."Notice the blameless version names system gaps (no validation, inadequate checklist) that you can actually fix. The blameful version names a person you can't "fix" and teaches everyone to hide mistakes.
Memory hook"second story." The blameful "first story" is who screwed up. The blameless "second story" asks why did this action make sense to a competent person at the time? The second story is where the real, fixable causes live. Replace "human error" with "what made the error easy?"
Root Cause Analysis: dig past the symptom
The point of a postmortem is the real root cause, not the surface trigger. Two techniques:
The 5 Whys — keep asking "why" until you hit something systemic:
Incident: customer data was exposed in an S3 bucket.
Why? The bucket was public.
Why? A developer set it public to debug and forgot to revert.
Why? There was no guardrail preventing public buckets.
Why? The org never enabled S3 Block Public Access at the account level.
Why? No policy/standard required it, and no detection flagged the drift.
→ Root causes: missing preventive control (account-level BPA) AND missing
detective control (no alert on public-bucket creation). Both are fixable.Memory hookthere is rarely ONE root causeMature postmortems find a chain of contributing factors — usually a missing preventive control AND a missing detective control AND a process gap. "Swiss cheese model": incidents happen when holes in multiple layers line up. List all the holes, not just the last one.
Contributing vs. root vs. trigger
the immediate event ("the config deployed at 14:03").
the systemic reason it was possible ("no validation gate existed").
things that made it worse or slower to catch ("alerting was too noisy, so the page was missed").
Anatomy of a Good Postmortem
1. TITLE & METADATA Incident ID, severity, date, authors, status
2. EXECUTIVE SUMMARY 3-5 sentences a VP can read: what happened, impact,
root cause, status. (Write this LAST, put it FIRST.)
3. IMPACT Quantified: users affected, data exposed, downtime,
$ cost, SLA/regulatory implications
4. TIMELINE Timestamped, factual sequence — detection → response →
containment → recovery. Include time-to-detect/respond.
5. ROOT CAUSE ANALYSIS The 5-Whys / contributing factors; preventive AND
detective gaps
6. WHAT WENT WELL Genuinely — fast detection, good runbook, clean comms.
(Reinforces good practices; keeps it balanced.)
7. WHAT WENT POORLY Blameless: the gaps, the confusion, the slow steps
8. ACTION ITEMS The most important section — see below
9. LESSONS LEARNED The durable takeaways for the org
10. APPENDICES IOCs, log excerpts, queries, referencesAction items make or break a postmortem
A postmortem with no owned, tracked action items is theatre. Each action item must be:
"Enable S3 Block Public Access org-wide," not "improve cloud security."
a named person/team, not "the team."
a ticket with a due date, reviewed to completion.
does it prevent recurrence, detect it faster, or mitigate impact?
Memory hookthe test of a postmortem is whether the same incident can happen tomorrow. If every action item shipped, would this recur? If yes, the analysis didn't reach the root cause or the actions are too weak. Action items are the only part that changes the future.
The Timeline — write it like an investigator
The timeline is the factual backbone. Rules:
incidents cross timezones; ambiguity is dangerous. 2026-06-09T14:03:12Z.
"14:03 config deployed (CloudTrail)" is fact; "the attacker likely pivoted here" is labelled analysis.
each entry references its source (log, alert, EDR event) so it's verifiable.
when you detected, who was paged, what you did. This is where MTTD and MTTR come from.
2026-06-09T13:58Z Initial access: phishing link clicked (proxy log)
2026-06-09T14:03Z Attacker authenticated to VPN with stolen creds (Okta log)
2026-06-09T14:20Z Lateral movement to fileserver (EDR process tree)
2026-06-09T15:10Z DETECTION: GuardDuty flags anomalous S3 access [MTTD ~72 min]
2026-06-09T15:14Z On-call paged; IR channel opened
2026-06-09T16:40Z CONTAINMENT: creds revoked, sessions killed, host isolated [MTTR ~90 min]
2026-06-09T18:00Z Eradication complete; monitoring for re-entryWriting Style for Each Audience
The same incident is written differently for different readers:
lead with business impact and risk, not packet captures. Plain language, no jargon. "Customer email addresses were exposed; no passwords or payment data. We've closed the gap and notified affected users."
technical depth, exact root cause, reproducible detail, specific fixes.
precise facts, what data, how many records, timeline relative to notification deadlines; careful, factual language (this may be discoverable).
honest, clear, non-alarmist; what happened, what data, what you're doing, what they should do.
Memory hookthe inverted pyramidLead with the conclusion (impact + status), then supporting detail, then deep appendices. A VP reads the first paragraph; an engineer reads to the appendix. Write so each can stop reading at the right depth and still be correctly informed. Write the executive summary last but place it first.
Severity Classification (so reports are comparable)
Define severities before incidents so everyone speaks the same language:
| Sev | Rough meaning | Example |
|---|---|---|
| SEV1 | Critical — active major impact | Ransomware spreading; customer data actively exfiltrating; production down |
| SEV2 | High — serious, contained or imminent | Confirmed intrusion, single system; credential compromise |
| SEV3 | Medium — limited impact | Isolated malware on one endpoint, no spread |
| SEV4 | Low — minimal | Blocked phishing attempt, policy violation |
Severity drives who's paged, how fast, and how formal the writeup. State your org's scale rather than guessing in an interview, but show you understand severity drives response.
Common Failure Modes (what makes reports bad)
names a culprit; kills future honesty.
nothing changes.
stops at the trigger ("the cert expired") not the system gap ("no cert-expiry monitoring").
"some users affected" is useless to decision-makers.
inaccurate; keep the live log.
the reader can't find what happened in three paragraphs of preamble.
omitting what went well loses lessons and morale.
memory fades, evidence ages, regulatory clocks tick.
Reusable Templates
SITREP (during the incident — keep it to one screen)
INCIDENT: <id> | SEV: <n> | STATUS: <investigating/contained/recovering>
AS OF: <UTC timestamp>
SITUATION: One line — what is happening.
IMPACT: What's affected right now (systems, users, data).
ACTIONS TAKEN: Bullet list since last SITREP.
NEXT STEPS: What we're doing next + ETA.
NEEDS: Decisions/resources required from leadership.Postmortem (after — the learning doc)
# Postmortem: <short title>
Incident ID | Severity | Date | Authors | Status: [Draft/Final]
## Executive Summary
<3-5 sentences: what happened, impact, root cause, current status>
## Impact
<Quantified: users, data, downtime, cost, regulatory>
## Timeline (UTC)
<timestamp> <fact, with evidence source>
...
Detection at <t> (MTTD: X) | Containment at <t> (MTTR: Y)
## Root Cause Analysis
<5 Whys; preventive gap AND detective gap; contributing factors>
## What Went Well
- ...
## What Went Poorly (blameless)
- ...
## Action Items
| # | Action | Owner | Type (prevent/detect/mitigate) | Ticket | Due |
|---|--------|-------|--------------------------------|--------|-----|
## Lessons Learned
- ...
## Appendices
IOCs, queries, log excerpts, referencesInterview Questions
A blameless postmortem assumes everyone acted reasonably given what they knew at the time and focuses on the systems and processes that allowed the incident, not on punishing individuals. It matters because it's the only way to get the truth: if people fear blame, they hide the details that reveal the real root cause, so you fix nothing and the incident recurs. Instead of "Alex pushed bad config," you write "a config change reached production because no validation gate caught this error class and the review checklist didn't cover it" — which names fixable system gaps. Blame drives the truth underground; blamelessness surfaces it, and the goal is learning, not accountability theatre.
It opens with metadata and an executive summary — three to five sentences a VP can read covering what happened, the impact, the root cause, and current status, written last but placed first. Then a quantified impact section (users, data, downtime, cost, regulatory implications), a factual UTC timeline citing evidence sources and including time-to-detect and time-to-respond, and the root cause analysis using something like the 5 Whys to reach systemic causes — ideally identifying both a missing preventive control and a missing detective one. Then a balanced "what went well" and a blameless "what went poorly," and the most important section: specific, owned, tracked action items categorised as prevent, detect, or mitigate. It closes with lessons learned and appendices of IOCs and queries. The test is whether the same incident could happen tomorrow if every action item shipped.
I use the 5 Whys, repeatedly asking why until I reach something systemic rather than the immediate trigger. If a bucket was public, "a developer set it public" isn't the root cause — keep going: why was there no guardrail preventing public buckets, why wasn't account-level Block Public Access enabled, why didn't anything detect the drift. That surfaces the real causes: a missing preventive control and a missing detective control. I also distinguish trigger, root cause, and contributing factors, and I assume there's rarely a single cause — using the Swiss cheese model, incidents happen when holes in several layers line up, so I document the whole chain of gaps, because fixing only the last one leaves the others open.
For executives I lead with business impact and risk in plain language — what data was affected, how many customers, whether it's contained, and what we're doing — and I keep it to a short summary with no packet-level detail. For engineers I provide the full technical depth: the exact root cause, reproducible detail, the timeline with evidence, and the specific fixes. It's the inverted pyramid — conclusion and impact first, supporting detail next, deep technical appendices last — so each audience can stop reading at the right depth and still be correctly informed. The facts are identical; the framing and altitude differ. For legal I'd be especially precise and factual since the document may be discoverable.
Because they're the only part that changes the future — the rest of the postmortem documents the past, but action items prevent recurrence. A postmortem with no owned, tracked actions is theatre. A good action item is specific ("enable S3 Block Public Access org-wide," not "improve cloud security"), owned by a named person or team, tracked as a ticket with a due date and followed to completion, and categorised by whether it prevents the incident, detects it faster, or mitigates its impact. The test of the whole exercise is simple: if every action item shipped, could this exact incident happen again tomorrow? If yes, either the root cause analysis fell short or the actions are too weak.