N · Checklists

Checklist: Assurance and Evaluation

Criteria before results, both error directions, and what a failed evaluation obliges.

The distinguishing question at S6 is not whether the system was tested but whether the criteria existed before the test did. Everything else on this list follows from that one discipline.

Depth scales with tier. A Tier 1 system needs the first set and little else; a Tier 4 system needs all of them.

Ticks are stored in this browser only. Nothing is sent anywhere, and clearing site data clears them.

A completed checklist is not evidence. The evidence is the record it told you to complete — the templates hold those.

Before the evaluation runs

  • Pass criteria written and dated, before any results exist System Owner
  • Negative criteria written: what the system must refuse, must not claim, must not disclose System Owner
  • Held-out set constructed from real inputs, including awkward ones ML engineering
  • Set documented well enough to be re-run and versioned ML engineering
  • Scope stated, including what will not be evaluated and why System Owner
  • Directional criteria set separately where errors are asymmetric Risk Lead

Coverage by tier

  • Tier 1–2: functional criteria, basic error rate, and the refusal behaviours
  • Tier 2+: directional error rates reported separately
  • Tier 3+: subgroup or stratum results with sample sizes that make them meaningful
  • Tier 3+: adversarial testing — injection through retrieved content, instruction override, exfiltration probes
  • Tier 3+: boundary probing for anything that acts, against the machine-readable boundary
  • Tier 4: reproducibility check — same input, same version, same output
  • Tier 4: degradation testing under volume, latency and partial failure

Handling results

  • Results reported per criterion, pass or fail, with the number
  • Failure analysis written: what the failures had in common
  • Omission graded explicitly for summarization and extraction tasks
  • External assurance referenced by version, never restated
  • Open issues recorded with an owner and a date
  • Results that contradict a recorded risk score routed back to classification

When the evaluation fails

  • The failure is recorded before any remediation begins
  • Remediation is a change, and the re-run scope is stated
  • Criteria are not relaxed to fit the result — if they were wrong, that is a separate, recorded decision
  • Repeat failures on the same criterion escalate rather than iterate
  • Where the system proceeds anyway, an explicit acceptance is recorded with a named accepter

Keeping assurance alive after deployment

  • Monitoring commitments derived from the criteria, not invented separately
  • Evaluation set retained and versioned for comparison over time
  • Re-run triggers agreed: model version, instruction version, corpus change, population change
  • Assurance re-run scheduled at least annually at Tier 3–4
  • LLM-based grading, if used, calibrated against human judgement on a sample