Checklist: Assurance and Evaluation
Criteria before results, both error directions, and what a failed evaluation obliges.
The distinguishing question at S6 is not whether the system was tested but whether the criteria existed before the test did. Everything else on this list follows from that one discipline.
Depth scales with tier. A Tier 1 system needs the first set and little else; a Tier 4 system needs all of them.
A completed checklist is not evidence. The evidence is the record it told you to complete — the templates hold those.
Before the evaluation runs
- Pass criteria written and dated, before any results exist System Owner
- Negative criteria written: what the system must refuse, must not claim, must not disclose System Owner
- Held-out set constructed from real inputs, including awkward ones ML engineering
- Set documented well enough to be re-run and versioned ML engineering
- Scope stated, including what will not be evaluated and why System Owner
- Directional criteria set separately where errors are asymmetric Risk Lead
Coverage by tier
- Tier 1–2: functional criteria, basic error rate, and the refusal behaviours
- Tier 2+: directional error rates reported separately
- Tier 3+: subgroup or stratum results with sample sizes that make them meaningful
- Tier 3+: adversarial testing — injection through retrieved content, instruction override, exfiltration probes
- Tier 3+: boundary probing for anything that acts, against the machine-readable boundary
- Tier 4: reproducibility check — same input, same version, same output
- Tier 4: degradation testing under volume, latency and partial failure
Handling results
- Results reported per criterion, pass or fail, with the number
- Failure analysis written: what the failures had in common
- Omission graded explicitly for summarization and extraction tasks
- External assurance referenced by version, never restated
- Open issues recorded with an owner and a date
- Results that contradict a recorded risk score routed back to classification
When the evaluation fails
- The failure is recorded before any remediation begins
- Remediation is a change, and the re-run scope is stated
- Criteria are not relaxed to fit the result — if they were wrong, that is a separate, recorded decision
- Repeat failures on the same criterion escalate rather than iterate
- Where the system proceeds anyway, an explicit acceptance is recorded with a named accepter
Keeping assurance alive after deployment
- Monitoring commitments derived from the criteria, not invented separately
- Evaluation set retained and versioned for comparison over time
- Re-run triggers agreed: model version, instruction version, corpus change, population change
- Assurance re-run scheduled at least annually at Tier 3–4
- LLM-based grading, if used, calibrated against human judgement on a sample