Practical Guide: Healthcare Delivery
Clinical safety management already exists and already works. The overlay's job is to connect it to the systems that were never treated as clinical, and to the ones that keep changing after clearance.
- Typical top tier
- Tier 4 — anything that alters triage order, treatment or discharge
- Already-owned ground
- Clinical safety case, incident reporting, device regulation, ethics review
- Hardest gate
- G3 — deployment authorization at the site, distinct from device clearance
- First record to fix
- The named clinical owner for every AI system already running
1. Where the risk actually concentrates
Healthcare arrives with a mature safety culture and a regulatory regime for devices, and both are narrower than the AI estate. A model cleared as a device is governed. An ambient documentation tool, a bed-management optimizer, a prior-authorization assistant and a patient-facing symptom checker may not be, and between them they touch more patients per day than the cleared device does.
The framework's contribution here is not clinical safety, which the sector does better than IRGF specifies. It is the chain of record from intent to runtime: which model version is running at this site, on this population, under whose authorization, with what evidence that it still performs the way it did at validation. Site-level performance is where the sector's real risk sits, because a model validated on one population and deployed on another is a different system in every respect that matters.
The second contribution is autonomy honesty. Clinical decision support is nominally advisory, and the framework's own scoring model flags advisory clinical systems as the case where formal autonomy most understates effective autonomy. A recommendation that is accepted 98% of the time, under time pressure, in a workflow that makes acceptance one click and rejection a form, is not advisory in any sense a patient would recognise.
Indicative classification of the systems this sector keeps building. The scores are illustrative, not authoritative: they show how the anchors in Chapter 6 read against sector facts. Score your own system; do not copy a row.
| Typical system | Illustrative D1–D5 | Tier | What is usually mis-scored |
|---|---|---|---|
| Deterioration, sepsis or risk-of-readmission prediction | D1 4 · D2 2–3 · D3 4 · D4 4 · D5 4 | Tier 4 | D2 read from the interface design rather than from observed acceptance rates |
| Imaging triage or worklist prioritization | D1 4 · D2 3 · D3 3–4 · D4 4 · D5 4 | Tier 4 | Reordering a worklist is treated as scheduling; a delayed finding is a clinical outcome |
| Ambient documentation and clinical note generation | D1 4 · D2 2 · D3 3 · D4 4 · D5 4 | Tier 3–4 | Scored as a productivity tool; the note becomes the record other clinicians and coders rely on |
| Symptom checker or patient-facing triage advice | D1 4 · D2 3–4 · D3 4 · D4 4 · D5 4 | Tier 4 | Advice given directly to a patient has no clinical reviewer at all |
| Prior authorization or coverage recommendation | D1 4 · D2 2–3 · D3 3 · D4 4 · D5 4 | Tier 3–4 | Delay is the harm and delay is invisible in accuracy metrics |
| Clinical coding and billing assistance | D1 3 · D2 2–3 · D3 3 · D4 4 · D5 4 | Tier 3 | Coding errors are financial until they are regulatory, and they propagate into the record |
| Bed, theatre and staffing optimization | D1 3–4 · D2 3 · D3 3 · D4 4 · D5 3 | Tier 3 | Treated as operations; the output determines who waits and for how long |
| Medicines interaction or dosing check | D1 4 · D2 2 · D3 4 · D4 4 · D5 4 | Tier 4 | Alert fatigue is a control failure, not a usability complaint |
| Research cohort discovery on de-identified data | D1 2 · D2 1 · D3 2 · D4 3 · D5 3–4 | Tier 2 | Re-identification risk in combination, scored as if de-identification were binary |
2. The regulatory interface
Regulatory note. The regimes below are named so each IRGF record can be pointed at the obligation it evidences, not to restate them. Applicability, thresholds and commencement dates differ by jurisdiction and several have moved during implementation. Nothing here is legal advice: confirm the current position with your own counsel, and record the answer in the Regulatory Overlay Reference so it is checkable later.
IRGF does not restate any of these obligations. It gives each one a record that carries the evidence, an owner, and a trigger that reopens it when the obligation or the system changes.
| Regime or standard | What it obliges in practice | IRGF record that carries the evidence |
|---|---|---|
| Medical device regulation where the software is a device (for example FDA 510(k)/De Novo, EU MDR) | Intended-use statement, clinical evaluation, quality system, post-market surveillance, and change control — with predetermined change control plans allowing some model updates without new submission. | Device clearance referenced from the AI Assurance Summary; any change outside the cleared change plan is a Major change in the AI Change Record |
| Clinical decision support transparency requirements (for example ONC HTI-1 source attributes for predictive decision support in certified health IT) | Disclosure of the intervention's purpose, inputs, development data, validation and maintenance to the clinicians relying on it. | Source attributes assembled from the Data Lineage and Sensitivity Record and the assurance summary; published as the derived AI System Card |
| Clinical safety standards for health IT (for example DCB0129 and DCB0160 in England, with a named Clinical Safety Officer) | A hazard log, a clinical safety case for manufacture and for deployment, and a named clinically qualified owner. | Hazard log referenced by the Control Matrix; the Clinical Safety Officer is the named accountable owner at G3 |
| Health information privacy and security law (for example HIPAA, GDPR special category data) | Minimum necessary access, audit of access to records, lawful basis for secondary use, and breach notification. | D5 sensitivity anchored at 4; access controls carried in the Control Matrix; secondary-use basis in the lineage record |
| Risk management and software lifecycle standards (ISO 14971, IEC 62304, and quality system requirements) | Hazard identification, risk control, traceability and lifecycle records. | Existing hazard and traceability records referenced, never duplicated. IRGF adds the runtime comparison these standards leave to the deployer |
| EU AI Act, where it applies | AI that is a safety component of a regulated device follows the Annex I route; some health-adjacent uses appear in Annex III. Obligations cover risk management, data quality, logging, human oversight and post-market monitoring. | Regulatory Overlay Reference, maintained per site as well as per system |
3. Calibrating the five dimensions
The dimensions do not change. What changes is what a 3 and a 4 look like when the subject matter is this sector, and which reading an assessor under delivery pressure reaches for first.
| Dimension | How to read it here | The mis-score to watch for |
|---|---|---|
| D1 Decision Consequence | Anything that changes what happens to a patient, when it happens, or what a clinician believes about them, is 4. That includes ordering effects: a worklist is a decision about who waits. | Scoring 3 because a clinician is nominally in the loop. D1 assumes the output is acted on. |
| D2 Autonomy | Score against measured acceptance, not workflow design. Where acceptance exceeds the agreed ceiling, effective autonomy is 3, whatever the interface says. | The advisory claim. It is the sector's most consequential mis-score and the framework says so in Chapter 6. |
| D3 Reversibility Deficit | A missed or delayed diagnosis, a dose given, and information disclosed to a patient are all 4. Very little in clinical care is reversible in the sense D3 means. | Scoring the record as reversible because it can be amended. |
| D4 Exposure and Scale | A system embedded in the electronic record reaches every encounter in every department that uses that workflow, including the ones nobody piloted it on. | Scoring the pilot ward. Also missing cascade: a wrong note feeds coding, research and the next clinician. |
| D5 Sensitivity and Uncertainty | Clinical data is 4. Generative summarization of clinical content carries uncertainty 3 or 4 because omission is the dangerous failure and omission is hard to detect. | Averaging sensitivity and uncertainty. Take the higher. |
4. One system, end to end
The overlay above is a map. This is one route across it: a single patient encounter followed from telemetry to outcome, with the record or control that attaches at each step.
5. What each gate adds
Additions only. Everything in the base gate definitions still applies; see the gate checklists for the common set.
| Gate | Sector addition | Why it is here |
|---|---|---|
| G1 | State the intended clinical use in the words a clinician would use, and state explicitly what the system is not for. Confirm whether the software is a device under the applicable regime. | Off-label drift starts with an intended-use statement written by the people funding the project. |
| G2 | Population comparison: the validation population against the deployment population, on the characteristics that matter clinically. Named data owner for every source, including the electronic record extract. | A model validated elsewhere is a hypothesis at your site until this comparison exists. |
| G3 | Site-level authorization by a named clinical owner, separate from device clearance. Agreed acceptance-rate ceiling and monitoring plan. Alert burden estimate for anything that interrupts. | Clearance says the device is safe in general. G3 says this deployment, on this population, with this workflow, is safe here. |
| G4 | Any model update outside the predetermined change control plan is Major. Any workflow change that alters where the clinician can intervene is an autonomy change. | The workflow is part of the system, and it is changed by people who never touch the model. |
| G5 | Records retention to the clinical standard, and a preserved ability to explain a historical recommendation for the duration of the liability window. | Clinical questions arrive years later, and they are about a version that no longer exists. |
6. Controls and evidence worth adding
| Control | Where it attaches | Evidence it produces |
|---|---|---|
| Named Clinical Safety Officer or equivalent as the accountable owner | Deployment Authorization Record | A named clinician, per site, with the authority to suspend |
| Acceptance-rate and override monitoring against a pre-agreed ceiling | Runtime signal; assurance cycle | Measured acceptance rate by unit and shift, with the re-score trigger stated |
| Subgroup performance monitoring on clinically relevant strata | Control Matrix; assurance cycle | Performance by stratum with the sample sizes that make it meaningful |
| Alert burden and fatigue measurement for interrupting systems | Runtime signal | Alerts per clinician per shift, and the dismissal rate trend |
| Documentation accuracy sampling for generated notes | Control Matrix | Sampled notes reviewed against source audio or encounter, with the omission rate recorded |
| Suspension runbook that a clinical lead can execute without engineering | Kill-switch record at Tier 3–4 | A tested procedure, with the last test date |
7. Runtime signals to wire first
Control Plane onboarding order matters more than coverage in the first year (Chapter 18). These are the signals that earn their place earliest in this sector.
| Signal | Drift category | Suggested response |
|---|---|---|
| Acceptance rate rises above the agreed ceiling | Behavioral, with a classification consequence | Re-score D2 and review the oversight design; the fix is usually the workflow, not the model |
| Deployment population diverges from the validation population | AI-model behavioral | Trigger the assurance re-run at Tier 3–4; a population shift is a silent performance change |
| Electronic record upgrade changes an extract, a code set or a field meaning | Data lineage and governance | Same-day owner alert; most clinical AI breakage originates here rather than in the model |
| Alert volume per clinician exceeds the agreed burden | Behavioral | Treat as a control failure and route to the clinical owner; fatigue disables the human control |
| Vendor pushes a model update to a hosted clinical service | AI-model behavioral | Verify it falls inside the cleared change plan; if not, it is a Major change and the deployment authorization is void until re-taken |
8. Failure modes this sector produces
Advisory in name
The tell. The system is scored D2 = 1, the interface has an accept button, and the accept rate is 97% at 3 a.m.
The response. Instrument acceptance before deployment and agree the ceiling at G2. Where acceptance is structurally certain, design the oversight point to require something the system did not supply — a reason code, a choice between genuinely different options.
The scribe that writes the record
The tell. Ambient documentation is procured as an efficiency measure by an operations team, and the generated note becomes the clinical record other decisions rely on.
The response. Score it on the decisions its output shapes. Sample for omission specifically: fluent notes that leave something out read better than accurate ones.
Cleared once, changed continuously
The tell. A cleared model is updated by the supplier under a managed service, and nobody at the site can say which version produced last month's recommendations.
The response. Version identity is a G3 requirement and a runtime signal. If the supplier cannot tell you which version ran on a given date, that is a procurement finding, not a technical one.
Governance stops at the department that bought it
The tell. A tool piloted in one unit spreads by word of mouth into three others with different populations.
The response. Population change is a Major change. Detect spread from access telemetry, not from the project plan.
Equity monitored once, at validation
The tell. Subgroup performance was assessed in the original study and never again at the site.
The response. Subgroup monitoring belongs in the assurance cycle with a named owner, sample sizes agreed in advance, and a stated action when a stratum degrades.
9. A ninety-day start
If the sector is yours and the framework is new, this is the order that produces something defensible fastest. It assumes one part-time architect and one risk lead, not a programme.
- Inventory every AI system touching a patient pathway, including features inside the electronic record and anything procured by a department rather than by IT.
- Name a clinical owner for each. Systems without one are the first backlog, ahead of any scoring work.
- Score the five highest-consequence systems with the clinical safety function in the room, and record the acceptance ceiling for each.
- Compare validation and deployment populations for the top two. Expect this to take longer than the scoring did.
- Instrument acceptance rate and alert burden on the two most-used systems.
- Write the suspension runbook for the highest-tier system and test it once, with the clinical lead executing it.
- Wire one drift signal: electronic-record extract change. It breaks more clinical AI than model change does.
- Run a G3 dry run on a system already live, and treat the gap list as the programme plan.