M · Guidelines

Guideline: Generative Systems and Retrieval

Prompts are configuration, grounding is architecture, and a demonstration is not an evaluation.

The failure this prevents

Generative systems break governance assumptions built for models with stable inputs and measurable accuracy. The behaviour changes when the prompt changes, when the corpus changes, when the supplier updates the model, and sometimes when nothing observable changes at all. The output is fluent whether or not it is right, which removes the informal signal that used to catch errors.

The result is a governance failure specific to this class: systems that pass review on a demonstration, change weekly through edits nobody classifies as change, and fail in ways their metrics do not measure.

Instruction versioned configuration Retrieval approved sources, staleness limits Model pinned version, change notification Post-check refusals, omissions, citations Action constrained surface, untrusted content
Five places a generative system changes behaviour, and only one of them is the model. Governance programmes that watch the model version alone miss four of the five — and the instruction and corpus columns are the two that change weekly.

[Practice recommendation] Everything on this page is advice from this documentation rather than a requirement of IRGF. Disagreeing with it does not put you outside the framework; all practice recommendations are collected in Appendix B.

Prompts and system instructions are configuration

A system instruction determines what the system will and will not do. Editing it is a configuration change to an approved system, and the framework's change model treats a substantive change to a system instruction as Material.

  • Version them. In the same repository discipline as code, with review, history and the ability to say what was running on a given date.
  • Bind the version to the authorization. The authorized configuration should name the instruction version, or drift detection has nothing to compare.
  • Separate the editable from the fixed. Tone and formatting can be freely edited; scope, refusals, boundaries and disclosure statements cannot. Make that split explicit or every edit becomes a governance question.
  • Test the boundary instructions specifically. The instruction that says what the system must refuse is the one worth an adversarial test.

Grounding is architecture, not content

Every retrieval source is a grounding source with an owner, a refresh cadence and a maximum staleness. This is where the two meanings of “data drift” separate: ungoverned change to grounding is an architecture and lineage concern, and distribution shift in model inputs is a model-performance concern. Conflating them routes alerts to people who cannot act on them.

Design question Why it is a governance question
What is in the corpus, and who approved each source?An unapproved source in a shared index is the most common confidentiality failure in retrieval systems.
What happens to a document with no classification?Fail closed. An ingestion pipeline that defaults unknown documents to “include” will eventually include the wrong thing.
How stale may a source be before the answer is untrustworthy?Confidently wrong answers from stale policy documents are worse than an unavailable system.
What does the system do when retrieval returns nothing relevant?The dangerous default is to answer anyway from parametric knowledge.
Is retrieval scoped by the requester's access rights?If not, the index has become an access-control bypass.
Which derived stores exist — indexes, embeddings, caches, evaluation sets?These are what deletion, retention and retirement have to reach.

Evaluation, not demonstration

A demonstration shows the system doing something. An evaluation states in advance what it must and must not do, measures both on a held-out set that represents deployment conditions, and records the failures.

  1. Write the criteria before the run, with the date. Include the negative criteria: what the system must refuse, must not claim, must not disclose.
  2. Build the set from real inputs, including the awkward ones. Sets built from imagined queries measure imagination.
  3. Grade for omission, not only error. In summarization and extraction the dangerous failure is what is missing, and fluent output hides it.
  4. Include adversarial cases at Tier 3–4: prompt injection through retrieved content, instruction override attempts, data exfiltration probes, boundary testing for anything that acts.
  5. Re-run on change — model version, instruction version, corpus change — and record what was not re-run and why.
  6. Keep the set stable so results are comparable over time, and version it when it changes.

Where an LLM grades the outputs, treat it as an instrument with its own error: calibrate it against human judgement on a sample, record the agreement rate, and do not let it grade the criteria it also generated.

Failure modes specific to this class

Injection through content. Instructions embedded in a retrieved document or an incoming email are executed by a system that cannot distinguish data from instruction. For any system that acts on retrieved content, treat every source as untrusted and constrain the action surface — this is an architecture control, not a prompt.

Silent scope expansion. A tool is added, a corpus is extended, an integration is enabled, and the system can now do something it was never authorized to do. This is autonomy change and it re-scores D2.

Confident omission. The output reads complete and is missing the material fact. Only sampling against the source catches it, and only if the sample checks for omission specifically.

Fabricated references. Citations, authorities and figures that are plausible and false. Verification against the source is the only control; asking the system to check itself is not one.

How to tell it is working

Every practice needs a failure signal, or it is a belief. These are the ones that show this guideline has stopped operating in your organization.

Signal What it means What to do
System instructions are edited without a change recordConfiguration change is happening outside the change processVersion the instructions and bind the version to the authorization
A source appears in the index that is not in the lineage recordIngestion is not enforcing the approved source listFail the build on unapproved sources; a preventive control here is cheap
Evaluation criteria were written after the resultsThe evaluation has become a demonstrationDate the criteria and require the date to precede the run
Only positive behaviours are testedRefusals, omissions and boundary cases are unmeasuredAdd negative criteria and omission grading to the set
The system answers when retrieval returns nothingThe fallback is parametric knowledge presented as groundedMake the no-retrieval path explicit and test it