Checklist: Data Readiness
What has to be true about the data before G2, and the two fields everyone leaves blank.
Data readiness is the gate most often failed for the same two reasons: a source with no owner, and no answer to how stale the data may be before the output stops being trustworthy.
Work source by source. Stopping at the first source with no owner is a legitimate outcome — that is the finding.
A completed checklist is not evidence. The evidence is the record it told you to complete — the templates hold those.
Sources
- Every source named, with its system of record, not the platform it arrives on Data Owner
- Each source has a named human owner Data Owner
- Basis for this specific use recorded — a basis for holding data is not a basis for this processing Risk Lead
- Contractual restrictions checked, including client or supplier terms that prohibit the processing Risk Lead
- Sources deliberately excluded are recorded, with the reason Architect
Classification and sensitivity
- Classification recorded per source
- Classification of the joined record assessed, where combination raises sensitivity
- Special categories identified explicitly, including inferences that amount to them
- Re-identification risk considered for anything described as anonymous or de-identified
- D5 sensitivity sub-score reconciled with what this record actually says
Freshness and grounding
- Refresh cadence recorded per source, with the mechanism
- Maximum staleness recorded — how old before the output is untrustworthy
- Action on breach of staleness stated: degrade, warn, or fail closed
- Behaviour when retrieval returns nothing relevant is defined and tested
- Retrieval scoped by the requester's access rights, where the corpus is not uniformly readable
- Unclassified documents fail closed at ingestion rather than defaulting to include
Derived stores and lifecycle
- Indexes, embeddings, caches and evaluation sets listed
- Deletion propagation to those stores tested at least once, on a real record
- Retention period set per source, against the longest applicable obligation
- Onward flows recorded: who consumes the output and what that makes them responsible for
- Rebuild path known: what it takes to reconstruct the index from approved sources
Quality and history
- Known quality issues recorded rather than left to be discovered by a model
- Historical decisions in training data assessed for practice you would not defend today
- Selection effects considered: does the data record what was examined rather than what was true?
- Population comparison done where a model was validated elsewhere
- Coverage gaps for smaller groups identified, with sample sizes stated
Blockers — do not pass G2 with any of these open
- A grounding source with no owner
- A source with no maximum staleness
- An ingestion path that accepts unclassified documents
- A derived store nobody has listed
- A processing purpose with no recorded basis