N · Checklists

Checklist: Risk Classification

Running a scoring session that produces the same tier whoever runs it, and a record that survives a challenge.

Classification is where implementations most often go quietly wrong, and the cause is almost never bad faith. It is anchors applied loosely, in a document, by one person who also owns the delivery date.

This checklist assumes forty minutes and two people. The full reasoning is in the scoring guideline.

Ticks are stored in this browser only. Nothing is sent anywhere, and clearing site data clears them.

A completed checklist is not evidence. The evidence is the record it told you to complete — the templates hold those.

Before the session

  • The use-case canvas exists and has been read by both scorers
  • The second scorer does not carry the delivery date for this system
  • The relevant sector overlay section 3 has been read (twenty minutes)
  • The decision the output supports is written in one sentence
  • Someone can answer what the system does when it is wrong, not only when it is right

D1 — Decision consequence severity

  • Scored on the consequence of being wrong, not the value of being right
  • Scored assuming the output is acted on — the human catch belongs to D2
  • Justification names who is harmed and how
  • The 3-versus-4 test applied: serious harm, or an inconvenience they could get fixed?
  • If D1 = 4, the safety-override floor is noted now, not discovered later

D2 — Autonomy

  • Scored against the highest-autonomy path the system has, including the routine bulk path
  • Straight-through, auto-close and auto-approve thresholds identified and counted as autonomy
  • Where a human reviews: the override rate is measured, or the floor is agreed today
  • Friction symmetry considered — is departing harder than accepting?
  • Planned autonomy trajectory recorded, so a later increase is a planned change

D3 — Reversibility deficit

  • Scored on consequence reversibility, never on mechanism reversibility
  • Asked explicitly: who has already seen, received or relied on the output?
  • Money movement, disclosure, communication and physical effect treated as irreversible
  • Appeal or complaint routes not counted as reversibility
  • Elapsed harm counted: a decision reversed after four months did not undo the four months

D4 — Exposure and scale

  • Counted the reach of the records the system writes into, not the login list
  • Cascade considered: which downstream systems or decisions consume this output
  • Scored on the deployment path, not the pilot population
  • External exposure identified, including indirect exposure through a partner

D5 — Sensitivity and uncertainty

  • Both sub-scores recorded separately
  • Composite taken as the higher of the two, not the average
  • Sensitivity assessed on the joined record, not source by source
  • Uncertainty scored honestly for generative components, where omission is the failure mode

Computing the tier

  • Impact = average(D1, D4, D5), rounded up — arithmetic shown in the record
  • Control Deficit = average(D2, D3), rounded up — arithmetic shown
  • Matrix cell recorded, not only the resulting tier
  • Override floor checked deliberately: D1 = 4, or D2 = 4 with D3 = 4
  • Result sanity-checked against a comparable system already classified

Countersignature and after

  • Countersigner was present, or scored independently and compared Risk Lead
  • Adjustments made by the countersigner are recorded, including none Risk Lead
  • Disagreements recorded with both positions and who decided Risk Lead
  • Effective-autonomy threshold and measurement method recorded System Owner
  • Re-score triggers noted, with the next scheduled review date Risk Lead
  • Previous scores preserved, not overwritten Risk Lead
  • Classification published where other teams can see it and argue with it AI Governance Lead

Re-score triggers — check quarterly that these are wired

  • Any change classified Material or Major
  • Any autonomy increase, of any size
  • User population change
  • New or changed grounding source at or above current sensitivity
  • Control Plane risk-relevant drift alert
  • Scheduled review: annual minimum at Tier 3–4
  • Override-rate evidence contradicting the recorded D2