Start incident notes as soon as an event needs coordinated investigation or could affect users, data, security, or a critical business process. The notes are a shared operational record: capture what was known at the time, what the team did, who made each decision, and which evidence supports it.
Use one time zone throughout—preferably UTC—and write facts before interpretations. Keep secrets, credentials, private keys, and unnecessary personal data out of the record. Link to access-controlled evidence instead of pasting sensitive log or payload content.
The useful way to frame the problem
Copy this header when opening an incident:
- Incident ID/title:
- Severity:
- Status: Investigating / Identified / Monitoring / Resolved
- Start time:
- Detection time:
- Recovery time:
- End time:
- Services affected:
- User/business impact:
- Current owner/incident lead:
- Trigger/detection source:
- Immediate mitigation:
- Links to logs/dashboards/tickets/runbooks:
Update status and ownership when they change. Leave unknown times marked Unknown rather than estimating them as fact; they can be reconstructed later from reliable evidence.
Capture the timeline
Add an entry whenever the observed state changes, somebody takes an action, the team makes a decision, or new evidence changes the working theory. Use precise timestamps and preserve the order in which information became available.
| Time | Observation/action | Owner | Evidence |
|---|---|---|---|
| Alert, user report, observed symptom, test, command, change, rollback, restore, or communication | Dashboard, log query, ticket, change record, screenshot, or runbook link | ||
Record both successful and unsuccessful actions. An unsuccessful mitigation can explain elapsed time and prevent another responder from repeating it. Do not paste authentication tokens or full sensitive payloads into the evidence column.
Separate impact from cause
During active response, write impact as an observed fact and cause as a labelled hypothesis. Do not turn an unverified explanation into a root-cause statement because it sounds plausible.
- Observed user/business impact: Which services, journeys, customers, regions, data, or internal processes were affected? What remained functional?
- Scope evidence: Which external checks, application errors, request samples, support reports, or business metrics establish that scope?
- Working hypothesis: What might explain the observations?
- Evidence for or against the hypothesis: Which result would confirm or reject it?
- Root cause (only after confirmed): State the causal mechanism and cite the confirming evidence. If confirmation is incomplete, say so.
- Contributing factors: Record conditions that increased impact, delayed detection, complicated recovery, or allowed recurrence without presenting them as the initiating cause.
Recovery does not prove cause. A restart may restore a service without establishing why it failed, and a rollback may reduce errors without proving that every changed line contributed. Finalise the cause only after the evidence supports it.
Record decisions
Record consequential decisions in the timeline and summarise them here:
- Decision: What did the team decide?
- Time and decision owner: Who had authority and when was the decision made?
- Evidence available: Which observations, risks, and unknowns informed it?
- Alternatives considered: What other response was considered?
- Expected result and risk: What should change, and what could the action make worse?
- Verification: How will the team know the action worked?
- Result: What actually happened?
Include decisions to deploy, stop a deployment, roll back, restore data, change DNS, rotate credentials, disable functionality, take a service offline, or communicate externally. For a restore, record the recovery point and any newer writes that may be lost. The VPS backup guide explains why a backup's existence is different from a verified recovery path.
List follow-up owners
Every follow-up needs a named owner, due date, and verification method. “Engineering will improve monitoring” is not an action item.
| Action | Owner | Due date | Verification |
|---|---|---|---|
| Test, alert simulation, restore evidence, reviewed change, or measured outcome | |||
Classify actions where useful: immediate risk reduction, permanent corrective work, detection improvement, recovery improvement, documentation, access, or customer follow-up. Track them in the team's normal work system and link the tickets from the incident record.
Turn notes into prevention
Complete this review after service is stable:
- What worked: Detection, ownership, communication, tooling, procedure, or recovery steps that reduced impact.
- What failed: Missing alerts, unclear ownership, unavailable access, unsafe changes, incomplete rollback, failed recovery, or misleading documentation.
- Detection gap: Could the team have detected the user-visible problem earlier or with a clearer signal?
- Recovery gap: Which dependency, access path, decision, or procedure delayed recovery?
- Recurrence control: Which change prevents the causal mechanism or limits its impact?
- Verification owner and date: Who proves each control works and when?
- Review participants: Who reviewed the final timeline, impact, cause, and actions?
- Finalised date: When was the record accepted as complete?
Keep the original timeline intact. Correct factual errors with an explicit amendment rather than silently rewriting the history responders relied on.