M · Guidelines

Guideline: Scoring Risk Consistently

How to run a scoring session so that two teams scoring the same system reach the same tier, and so that the tier survives contact with a delivery date.

The failure this prevents

Nearly every tier-scaled decision in the framework depends on classification being right, and classification is where implementations most often go quietly wrong. Not through bad faith: through anchors applied loosely by people who also own the delivery date, in a document rather than a conversation, with no independent reader.

The characteristic outcome is a system two tiers below where it belongs, with five individually arguable scores. Chapter 7, Example D shows exactly this: a Tier 4 reconciliation agent scored into Tier 2 by reading reversibility off the mechanism and exposure off the user count.

Prepare one sentence: what happens after the output Score justification first, then the number Compute both axes, rounding shown, floor checked Countersign independent, adjustments recorded Arm triggers thresholds agreed before they are known
The scoring session. The last step is the one most often skipped and the one that determines whether the classification survives: thresholds agreed while nobody knows the answer are policy, and the same thresholds agreed afterwards are negotiation.

[Practice recommendation] Everything on this page is advice from this documentation rather than a requirement of IRGF. Disagreeing with it does not put you outside the framework; all practice recommendations are collected in Appendix B.

Score in a session, not in a document

A scoring session takes forty minutes for a first system and twenty thereafter. Two people minimum. One of them must not carry the delivery date — that is the whole control, and it is cheaper than any audit.

  1. Read the decision aloud. One sentence: what happens in the world after this output appears. Most scoring errors are visible in this sentence before any number is written.
  2. Write the justification before the score. The sentence corrects the number more often than the number corrects the sentence.
  3. Take the dimensions in order — D1, D2, D3, D4, D5 — because each frames the next. D1 assumes the output is acted on; D2 is where the human catch belongs.
  4. Compute both axes out loud, including the rounding. Show the arithmetic in the record.
  5. Check the override floor last, deliberately. It applies without judgement, which means it is easy to forget.
  6. Record the disagreement, not just the outcome. A record that hides which score was disputed loses its most useful content.

The four anchor questions that resolve most disputes

When the argument is about Ask this Why it settles it
D1, and someone is arguing for 3 over 4Would a reasonable person affected by a wrong output call this a serious harm, or an inconvenience they could get fixed?It moves the frame from the organization's loss to the affected person's, which is what D1 asks about.
D2, and “a human approves each one”What proportion does that human change or reject, and how do we know?Formal autonomy is a design fact; effective autonomy is a measurement. Where nobody has measured, the honest score is the higher one until instrumentation exists.
D3, and “we can roll it back”Roll back what — the system state or the consequence? Who has already seen, received or relied on it?D3 is scored on consequence reversibility. Communication, disclosure and money movement do not reverse.
D4, and “only our team uses it”What records does it write into, and who reads those?Reach is a property of the output's destination, not the login list.

The countersignature, and making it real

The countersignature is the framework's main defence against tier gaming, and it fails in a specific way: a countersigner who reads the completed record and signs it. That is review of a document, not of a classification.

  • The countersigner attends the session, or scores independently and compares. Reading afterwards is the weakest available option.
  • They own two dimensions to challenge by default: D3 and D4, because those carry the most common errors.
  • Their adjustments are recorded. A countersigner whose adjustments are always zero is not operating, and that pattern is measurable across a quarter.
  • At Tier 3–4 the countersigner is outside the delivery line entirely.

What the countersignature does not catch

The framework states this openly: the override floor catches extremes, and the mid-range is where gaming survives. A system scored D1 = 3 that arguably warrants 4 lands two tiers lower with no automatic backstop. Three things help, none of which is a solution.

Thresholds agreed before the number is known. Override-rate floors, acceptance ceilings and error-rate limits set at G2 are arguments about a policy; set at the first assurance cycle they are arguments about a result.

Cross-system comparison. Once a quarter, list every system at Tier 2 with D1 = 3 and read them together. Inconsistency between systems is far easier to see than error within one.

Score distribution monitoring. If the portfolio's scores cluster just below every threshold that triggers work, that is a finding about the process, not about the systems.

Re-scoring is normal

Six triggers force a re-score: any Material or Major change; any autonomy increase; a user population change; a new or changed grounding source at or above current sensitivity; a risk-relevant drift alert; and the scheduled review, annual as a minimum at Tier 3–4. Add a seventh from practice: override-rate evidence that contradicts the recorded D2.

Preserve history rather than overwriting. The previous score with its date and the reason it changed is evidence of a functioning process; a record showing only the current tier is evidence of nothing.

How to tell it is working

Every practice needs a failure signal, or it is a belief. These are the ones that show this guideline has stopped operating in your organization.

Signal What it means What to do
Countersigner adjustments are consistently zeroThe control is being performed as a signature rather than a reviewMove the countersigner into the session, or rotate the role
Scores cluster just below thresholds that trigger extra workThe tier is being reverse-engineered from the desired processRun the cross-system comparison and re-score the outliers with a different pair
Re-scores almost never change a tierRe-scoring has become a formality applied to the same recordCheck the triggers are actually firing; unfired triggers usually mean nobody wired the signal
Two systems with the same shape sit at different tiersAnchors are being applied inconsistently between teamsPublish both classifications side by side and let the difference be argued in public
Scoring is done by one person in a documentThe session control has quietly lapsedReinstate it; this single change recovers most classification quality