Guideline: Scoring Risk Consistently
How to run a scoring session so that two teams scoring the same system reach the same tier, and so that the tier survives contact with a delivery date.
The failure this prevents
Nearly every tier-scaled decision in the framework depends on classification being right, and classification is where implementations most often go quietly wrong. Not through bad faith: through anchors applied loosely by people who also own the delivery date, in a document rather than a conversation, with no independent reader.
The characteristic outcome is a system two tiers below where it belongs, with five individually arguable scores. Chapter 7, Example D shows exactly this: a Tier 4 reconciliation agent scored into Tier 2 by reading reversibility off the mechanism and exposure off the user count.
[Practice recommendation] Everything on this page is advice from this documentation rather than a requirement of IRGF. Disagreeing with it does not put you outside the framework; all practice recommendations are collected in Appendix B.
Score in a session, not in a document
A scoring session takes forty minutes for a first system and twenty thereafter. Two people minimum. One of them must not carry the delivery date — that is the whole control, and it is cheaper than any audit.
- Read the decision aloud. One sentence: what happens in the world after this output appears. Most scoring errors are visible in this sentence before any number is written.
- Write the justification before the score. The sentence corrects the number more often than the number corrects the sentence.
- Take the dimensions in order — D1, D2, D3, D4, D5 — because each frames the next. D1 assumes the output is acted on; D2 is where the human catch belongs.
- Compute both axes out loud, including the rounding. Show the arithmetic in the record.
- Check the override floor last, deliberately. It applies without judgement, which means it is easy to forget.
- Record the disagreement, not just the outcome. A record that hides which score was disputed loses its most useful content.
The four anchor questions that resolve most disputes
| When the argument is about | Ask this | Why it settles it |
|---|---|---|
| D1, and someone is arguing for 3 over 4 | Would a reasonable person affected by a wrong output call this a serious harm, or an inconvenience they could get fixed? | It moves the frame from the organization's loss to the affected person's, which is what D1 asks about. |
| D2, and “a human approves each one” | What proportion does that human change or reject, and how do we know? | Formal autonomy is a design fact; effective autonomy is a measurement. Where nobody has measured, the honest score is the higher one until instrumentation exists. |
| D3, and “we can roll it back” | Roll back what — the system state or the consequence? Who has already seen, received or relied on it? | D3 is scored on consequence reversibility. Communication, disclosure and money movement do not reverse. |
| D4, and “only our team uses it” | What records does it write into, and who reads those? | Reach is a property of the output's destination, not the login list. |
The countersignature, and making it real
The countersignature is the framework's main defence against tier gaming, and it fails in a specific way: a countersigner who reads the completed record and signs it. That is review of a document, not of a classification.
- The countersigner attends the session, or scores independently and compares. Reading afterwards is the weakest available option.
- They own two dimensions to challenge by default: D3 and D4, because those carry the most common errors.
- Their adjustments are recorded. A countersigner whose adjustments are always zero is not operating, and that pattern is measurable across a quarter.
- At Tier 3–4 the countersigner is outside the delivery line entirely.
What the countersignature does not catch
The framework states this openly: the override floor catches extremes, and the mid-range is where gaming survives. A system scored D1 = 3 that arguably warrants 4 lands two tiers lower with no automatic backstop. Three things help, none of which is a solution.
Thresholds agreed before the number is known. Override-rate floors, acceptance ceilings and error-rate limits set at G2 are arguments about a policy; set at the first assurance cycle they are arguments about a result.
Cross-system comparison. Once a quarter, list every system at Tier 2 with D1 = 3 and read them together. Inconsistency between systems is far easier to see than error within one.
Score distribution monitoring. If the portfolio's scores cluster just below every threshold that triggers work, that is a finding about the process, not about the systems.
Re-scoring is normal
Six triggers force a re-score: any Material or Major change; any autonomy increase; a user population change; a new or changed grounding source at or above current sensitivity; a risk-relevant drift alert; and the scheduled review, annual as a minimum at Tier 3–4. Add a seventh from practice: override-rate evidence that contradicts the recorded D2.
Preserve history rather than overwriting. The previous score with its date and the reason it changed is evidence of a functioning process; a record showing only the current tier is evidence of nothing.
How to tell it is working
Every practice needs a failure signal, or it is a belief. These are the ones that show this guideline has stopped operating in your organization.
| Signal | What it means | What to do |
|---|---|---|
| Countersigner adjustments are consistently zero | The control is being performed as a signature rather than a review | Move the countersigner into the session, or rotate the role |
| Scores cluster just below thresholds that trigger extra work | The tier is being reverse-engineered from the desired process | Run the cross-system comparison and re-score the outliers with a different pair |
| Re-scores almost never change a tier | Re-scoring has become a formality applied to the same record | Check the triggers are actually firing; unfired triggers usually mean nobody wired the signal |
| Two systems with the same shape sit at different tiers | Anchors are being applied inconsistently between teams | Publish both classifications side by side and let the difference be argued in public |
| Scoring is done by one person in a document | The session control has quietly lapsed | Reinstate it; this single change recovers most classification quality |