Guideline: Oversight That Is Real
Automation bias is well established, review under volume is not review, and “human in the loop” is a claim until something measures it.
The failure this prevents
This is the framework's most important open weakness and it says so directly. A system scored D2 = 1 because a human makes every decision may carry far higher effective autonomy if that human approves essentially everything. Stress-testing surfaced it most sharply for advisory clinical and advisory financial systems, where the formal score understated real autonomy substantially.
IRGF does not solve automation bias. What it asks you to do is detect it, and then either fix the design or fix the score. Both are legitimate; neither is optional.
[Practice recommendation] Everything on this page is advice from this documentation rather than a requirement of IRGF. Disagreeing with it does not put you outside the framework; all practice recommendations are collected in Appendix B.
Measure first: the override rate
For any system scored D2 = 1 or 2 where D1 is 3 or 4, instrument the proportion of recommendations the human changes or rejects, and review it at the first post-deployment assurance cycle.
- Set the floor in advance. Before deployment, at G2 or G3, agree the override rate below which the recorded D2 is presumed wrong. Any number agreed after the measurement exists becomes a negotiation.
- Measure over meaningful volume. A 0% override rate over twelve decisions says nothing; over five hundred it says the reviewer is not reviewing.
- Segment it. By shift, by reviewer, by workload. Oversight fails under load, and an average hides exactly the condition you care about.
- Breach triggers a re-score, not a note. The system's recorded tier is wrong; that is a classification event with consequences.
Two adjacent signals are worth collecting where override rate is unavailable: time per decision (compared against the time the review actually needs) and reason-code diversity (a reviewer selecting the same code every time is not deciding).
Two design moves that beat any policy
Make the human supply something the system did not. A reason code chosen from meaningfully different options, a selection between genuinely distinct alternatives, a piece of context the system cannot see. If the only available action is approve, the interface has already decided.
Cap autonomy in architecture, not in policy. A system with no technical path to act cannot drift into acting. A system that can act but is instructed not to will, eventually, under pressure, with the best of intentions.
Symmetric friction
The most common structural cause of nominal oversight is asymmetry: accepting takes one click, departing requires a written justification and a manager's approval. Whatever the policy says, the system is now the decision-maker and the human is an audit trail.
Symmetry does not mean making acceptance harder. Where a genuine reason exists for extra scrutiny on departures, record that the oversight is asymmetric and score D2 accordingly, rather than claiming a review that the workflow discourages.
Oversight has a capacity, and it is smaller than you think
| Question | Why it decides whether oversight exists |
|---|---|
| How many decisions per reviewer per hour? | Above a few dozen, meaningful review of anything but exceptions is implausible. Say so in the record rather than discovering it in an incident. |
| What does the reviewer see besides the recommendation? | A recommendation without its grounds can only be accepted or refused arbitrarily. |
| What happens at peak volume? | Oversight degrades exactly when consequences concentrate. Design the degradation deliberately: queue, defer, or fail closed. |
| Who trained the reviewers on the failure modes? | Reviewers who do not know how the system fails cannot catch the failures. |
| What is the escalation route when unsure? | If the only options are accept and reject, uncertainty resolves as acceptance. |
When real oversight is not achievable
Sometimes volume, latency or expertise make genuine review impossible. That is an acceptable finding and an unacceptable pretence. The honest responses are: raise D2 to match reality and accept the higher tier; restrict the system's scope so the automated path carries less consequence; or add a preventive constraint so the system cannot produce the outcome that needed catching.
Sampling is a fourth option with a caveat: a human reviewing 2% of decisions is a quality control, not an oversight point, and D2 should be scored as if the other 98% were unreviewed — because they are.
How to tell it is working
Every practice needs a failure signal, or it is a belief. These are the ones that show this guideline has stopped operating in your organization.
| Signal | What it means | What to do |
|---|---|---|
| Override rate below the agreed floor over meaningful volume | Effective autonomy exceeds the recorded score | Re-score D2 and review the oversight design; the fix is usually the workflow |
| Time per decision is far below the time the review needs | The oversight point is nominal | Redesign the interaction or restrict the automated path |
| Reviewers select the same reason code almost always | The reason code has become a formality | Reduce the option list to meaningfully different choices, or remove it and stop claiming it as a control |
| Departures require more effort than acceptances | The workflow has made the recommendation into the decision | Equalise the friction or record the asymmetry and score accordingly |
| Oversight metrics exist but nobody has a threshold | Measurement without a pre-agreed action is reporting | Set the floor now, in writing, before the next review cycle |