Checklist: Risk Classification
Running a scoring session that produces the same tier whoever runs it, and a record that survives a challenge.
Classification is where implementations most often go quietly wrong, and the cause is almost never bad faith. It is anchors applied loosely, in a document, by one person who also owns the delivery date.
This checklist assumes forty minutes and two people. The full reasoning is in the scoring guideline.
A completed checklist is not evidence. The evidence is the record it told you to complete — the templates hold those.
Before the session
- The use-case canvas exists and has been read by both scorers
- The second scorer does not carry the delivery date for this system
- The relevant sector overlay section 3 has been read (twenty minutes)
- The decision the output supports is written in one sentence
- Someone can answer what the system does when it is wrong, not only when it is right
D1 — Decision consequence severity
- Scored on the consequence of being wrong, not the value of being right
- Scored assuming the output is acted on — the human catch belongs to D2
- Justification names who is harmed and how
- The 3-versus-4 test applied: serious harm, or an inconvenience they could get fixed?
- If D1 = 4, the safety-override floor is noted now, not discovered later
D2 — Autonomy
- Scored against the highest-autonomy path the system has, including the routine bulk path
- Straight-through, auto-close and auto-approve thresholds identified and counted as autonomy
- Where a human reviews: the override rate is measured, or the floor is agreed today
- Friction symmetry considered — is departing harder than accepting?
- Planned autonomy trajectory recorded, so a later increase is a planned change
D3 — Reversibility deficit
- Scored on consequence reversibility, never on mechanism reversibility
- Asked explicitly: who has already seen, received or relied on the output?
- Money movement, disclosure, communication and physical effect treated as irreversible
- Appeal or complaint routes not counted as reversibility
- Elapsed harm counted: a decision reversed after four months did not undo the four months
D4 — Exposure and scale
- Counted the reach of the records the system writes into, not the login list
- Cascade considered: which downstream systems or decisions consume this output
- Scored on the deployment path, not the pilot population
- External exposure identified, including indirect exposure through a partner
D5 — Sensitivity and uncertainty
- Both sub-scores recorded separately
- Composite taken as the higher of the two, not the average
- Sensitivity assessed on the joined record, not source by source
- Uncertainty scored honestly for generative components, where omission is the failure mode
Computing the tier
- Impact = average(D1, D4, D5), rounded up — arithmetic shown in the record
- Control Deficit = average(D2, D3), rounded up — arithmetic shown
- Matrix cell recorded, not only the resulting tier
- Override floor checked deliberately: D1 = 4, or D2 = 4 with D3 = 4
- Result sanity-checked against a comparable system already classified
Countersignature and after
- Countersigner was present, or scored independently and compared Risk Lead
- Adjustments made by the countersigner are recorded, including none Risk Lead
- Disagreements recorded with both positions and who decided Risk Lead
- Effective-autonomy threshold and measurement method recorded System Owner
- Re-score triggers noted, with the next scheduled review date Risk Lead
- Previous scores preserved, not overwritten Risk Lead
- Classification published where other teams can see it and argue with it AI Governance Lead
Re-score triggers — check quarterly that these are wired
- Any change classified Material or Major
- Any autonomy increase, of any size
- User population change
- New or changed grounding source at or above current sensitivity
- Control Plane risk-relevant drift alert
- Scheduled review: annual minimum at Tier 3–4
- Override-rate evidence contradicting the recorded D2