Guideline: Generative Systems and Retrieval
Prompts are configuration, grounding is architecture, and a demonstration is not an evaluation.
The failure this prevents
Generative systems break governance assumptions built for models with stable inputs and measurable accuracy. The behaviour changes when the prompt changes, when the corpus changes, when the supplier updates the model, and sometimes when nothing observable changes at all. The output is fluent whether or not it is right, which removes the informal signal that used to catch errors.
The result is a governance failure specific to this class: systems that pass review on a demonstration, change weekly through edits nobody classifies as change, and fail in ways their metrics do not measure.
[Practice recommendation] Everything on this page is advice from this documentation rather than a requirement of IRGF. Disagreeing with it does not put you outside the framework; all practice recommendations are collected in Appendix B.
Prompts and system instructions are configuration
A system instruction determines what the system will and will not do. Editing it is a configuration change to an approved system, and the framework's change model treats a substantive change to a system instruction as Material.
- Version them. In the same repository discipline as code, with review, history and the ability to say what was running on a given date.
- Bind the version to the authorization. The authorized configuration should name the instruction version, or drift detection has nothing to compare.
- Separate the editable from the fixed. Tone and formatting can be freely edited; scope, refusals, boundaries and disclosure statements cannot. Make that split explicit or every edit becomes a governance question.
- Test the boundary instructions specifically. The instruction that says what the system must refuse is the one worth an adversarial test.
Grounding is architecture, not content
Every retrieval source is a grounding source with an owner, a refresh cadence and a maximum staleness. This is where the two meanings of “data drift” separate: ungoverned change to grounding is an architecture and lineage concern, and distribution shift in model inputs is a model-performance concern. Conflating them routes alerts to people who cannot act on them.
| Design question | Why it is a governance question |
|---|---|
| What is in the corpus, and who approved each source? | An unapproved source in a shared index is the most common confidentiality failure in retrieval systems. |
| What happens to a document with no classification? | Fail closed. An ingestion pipeline that defaults unknown documents to “include” will eventually include the wrong thing. |
| How stale may a source be before the answer is untrustworthy? | Confidently wrong answers from stale policy documents are worse than an unavailable system. |
| What does the system do when retrieval returns nothing relevant? | The dangerous default is to answer anyway from parametric knowledge. |
| Is retrieval scoped by the requester's access rights? | If not, the index has become an access-control bypass. |
| Which derived stores exist — indexes, embeddings, caches, evaluation sets? | These are what deletion, retention and retirement have to reach. |
Evaluation, not demonstration
A demonstration shows the system doing something. An evaluation states in advance what it must and must not do, measures both on a held-out set that represents deployment conditions, and records the failures.
- Write the criteria before the run, with the date. Include the negative criteria: what the system must refuse, must not claim, must not disclose.
- Build the set from real inputs, including the awkward ones. Sets built from imagined queries measure imagination.
- Grade for omission, not only error. In summarization and extraction the dangerous failure is what is missing, and fluent output hides it.
- Include adversarial cases at Tier 3–4: prompt injection through retrieved content, instruction override attempts, data exfiltration probes, boundary testing for anything that acts.
- Re-run on change — model version, instruction version, corpus change — and record what was not re-run and why.
- Keep the set stable so results are comparable over time, and version it when it changes.
Where an LLM grades the outputs, treat it as an instrument with its own error: calibrate it against human judgement on a sample, record the agreement rate, and do not let it grade the criteria it also generated.
Failure modes specific to this class
Injection through content. Instructions embedded in a retrieved document or an incoming email are executed by a system that cannot distinguish data from instruction. For any system that acts on retrieved content, treat every source as untrusted and constrain the action surface — this is an architecture control, not a prompt.
Silent scope expansion. A tool is added, a corpus is extended, an integration is enabled, and the system can now do something it was never authorized to do. This is autonomy change and it re-scores D2.
Confident omission. The output reads complete and is missing the material fact. Only sampling against the source catches it, and only if the sample checks for omission specifically.
Fabricated references. Citations, authorities and figures that are plausible and false. Verification against the source is the only control; asking the system to check itself is not one.
How to tell it is working
Every practice needs a failure signal, or it is a belief. These are the ones that show this guideline has stopped operating in your organization.
| Signal | What it means | What to do |
|---|---|---|
| System instructions are edited without a change record | Configuration change is happening outside the change process | Version the instructions and bind the version to the authorization |
| A source appears in the index that is not in the lineage record | Ingestion is not enforcing the approved source list | Fail the build on unapproved sources; a preventive control here is cheap |
| Evaluation criteria were written after the results | The evaluation has become a demonstration | Date the criteria and require the date to precede the run |
| Only positive behaviours are tested | Refusals, omissions and boundary cases are unmeasured | Add negative criteria and omission grading to the set |
| The system answers when retrieval returns nothing | The fallback is parametric knowledge presented as grounded | Make the no-retrieval path explicit and test it |