K · Practical guide

Practical Guide: Healthcare Delivery

Clinical safety management already exists and already works. The overlay's job is to connect it to the systems that were never treated as clinical, and to the ones that keep changing after clearance.

Typical top tier
Tier 4 — anything that alters triage order, treatment or discharge
Already-owned ground
Clinical safety case, incident reporting, device regulation, ethics review
Hardest gate
G3 — deployment authorization at the site, distinct from device clearance
First record to fix
The named clinical owner for every AI system already running

1. Where the risk actually concentrates

Healthcare arrives with a mature safety culture and a regulatory regime for devices, and both are narrower than the AI estate. A model cleared as a device is governed. An ambient documentation tool, a bed-management optimizer, a prior-authorization assistant and a patient-facing symptom checker may not be, and between them they touch more patients per day than the cleared device does.

The framework's contribution here is not clinical safety, which the sector does better than IRGF specifies. It is the chain of record from intent to runtime: which model version is running at this site, on this population, under whose authorization, with what evidence that it still performs the way it did at validation. Site-level performance is where the sector's real risk sits, because a model validated on one population and deployed on another is a different system in every respect that matters.

The second contribution is autonomy honesty. Clinical decision support is nominally advisory, and the framework's own scoring model flags advisory clinical systems as the case where formal autonomy most understates effective autonomy. A recommendation that is accepted 98% of the time, under time pressure, in a workflow that makes acceptance one click and rejection a form, is not advisory in any sense a patient would recognise.

Access triage, appointments, referral Tier 3 Diagnosis imaging, deterioration prediction Tier 4 Treatment dosing, pathway and discharge support Tier 4 Record ambient notes, coding Tier 3 Operations beds, theatres, rostering Tier 3
Nothing in a patient pathway is low tier. The two columns hospitals govern least — the generated record and operational optimization — sit one step below diagnosis, because a note other clinicians rely on and a decision about who waits are both clinical acts.

Indicative classification of the systems this sector keeps building. The scores are illustrative, not authoritative: they show how the anchors in Chapter 6 read against sector facts. Score your own system; do not copy a row.

Typical system Illustrative D1–D5 Tier What is usually mis-scored
Deterioration, sepsis or risk-of-readmission predictionD1 4 · D2 2–3 · D3 4 · D4 4 · D5 4Tier 4D2 read from the interface design rather than from observed acceptance rates
Imaging triage or worklist prioritizationD1 4 · D2 3 · D3 3–4 · D4 4 · D5 4Tier 4Reordering a worklist is treated as scheduling; a delayed finding is a clinical outcome
Ambient documentation and clinical note generationD1 4 · D2 2 · D3 3 · D4 4 · D5 4Tier 3–4Scored as a productivity tool; the note becomes the record other clinicians and coders rely on
Symptom checker or patient-facing triage adviceD1 4 · D2 3–4 · D3 4 · D4 4 · D5 4Tier 4Advice given directly to a patient has no clinical reviewer at all
Prior authorization or coverage recommendationD1 4 · D2 2–3 · D3 3 · D4 4 · D5 4Tier 3–4Delay is the harm and delay is invisible in accuracy metrics
Clinical coding and billing assistanceD1 3 · D2 2–3 · D3 3 · D4 4 · D5 4Tier 3Coding errors are financial until they are regulatory, and they propagate into the record
Bed, theatre and staffing optimizationD1 3–4 · D2 3 · D3 3 · D4 4 · D5 3Tier 3Treated as operations; the output determines who waits and for how long
Medicines interaction or dosing checkD1 4 · D2 2 · D3 4 · D4 4 · D5 4Tier 4Alert fatigue is a control failure, not a usability complaint
Research cohort discovery on de-identified dataD1 2 · D2 1 · D3 2 · D4 3 · D5 3–4Tier 2Re-identification risk in combination, scored as if de-identification were binary

2. The regulatory interface

Regulatory note. The regimes below are named so each IRGF record can be pointed at the obligation it evidences, not to restate them. Applicability, thresholds and commencement dates differ by jurisdiction and several have moved during implementation. Nothing here is legal advice: confirm the current position with your own counsel, and record the answer in the Regulatory Overlay Reference so it is checkable later.

IRGF does not restate any of these obligations. It gives each one a record that carries the evidence, an owner, and a trigger that reopens it when the obligation or the system changes.

Regime or standard What it obliges in practice IRGF record that carries the evidence
Medical device regulation where the software is a device (for example FDA 510(k)/De Novo, EU MDR)Intended-use statement, clinical evaluation, quality system, post-market surveillance, and change control — with predetermined change control plans allowing some model updates without new submission.Device clearance referenced from the AI Assurance Summary; any change outside the cleared change plan is a Major change in the AI Change Record
Clinical decision support transparency requirements (for example ONC HTI-1 source attributes for predictive decision support in certified health IT)Disclosure of the intervention's purpose, inputs, development data, validation and maintenance to the clinicians relying on it.Source attributes assembled from the Data Lineage and Sensitivity Record and the assurance summary; published as the derived AI System Card
Clinical safety standards for health IT (for example DCB0129 and DCB0160 in England, with a named Clinical Safety Officer)A hazard log, a clinical safety case for manufacture and for deployment, and a named clinically qualified owner.Hazard log referenced by the Control Matrix; the Clinical Safety Officer is the named accountable owner at G3
Health information privacy and security law (for example HIPAA, GDPR special category data)Minimum necessary access, audit of access to records, lawful basis for secondary use, and breach notification.D5 sensitivity anchored at 4; access controls carried in the Control Matrix; secondary-use basis in the lineage record
Risk management and software lifecycle standards (ISO 14971, IEC 62304, and quality system requirements)Hazard identification, risk control, traceability and lifecycle records.Existing hazard and traceability records referenced, never duplicated. IRGF adds the runtime comparison these standards leave to the deployer
EU AI Act, where it appliesAI that is a safety component of a regulated device follows the Annex I route; some health-adjacent uses appear in Annex III. Obligations cover risk management, data quality, logging, human oversight and post-market monitoring.Regulatory Overlay Reference, maintained per site as well as per system

3. Calibrating the five dimensions

The dimensions do not change. What changes is what a 3 and a 4 look like when the subject matter is this sector, and which reading an assessor under delivery pressure reaches for first.

Dimension How to read it here The mis-score to watch for
D1 Decision ConsequenceAnything that changes what happens to a patient, when it happens, or what a clinician believes about them, is 4. That includes ordering effects: a worklist is a decision about who waits.Scoring 3 because a clinician is nominally in the loop. D1 assumes the output is acted on.
D2 AutonomyScore against measured acceptance, not workflow design. Where acceptance exceeds the agreed ceiling, effective autonomy is 3, whatever the interface says.The advisory claim. It is the sector's most consequential mis-score and the framework says so in Chapter 6.
D3 Reversibility DeficitA missed or delayed diagnosis, a dose given, and information disclosed to a patient are all 4. Very little in clinical care is reversible in the sense D3 means.Scoring the record as reversible because it can be amended.
D4 Exposure and ScaleA system embedded in the electronic record reaches every encounter in every department that uses that workflow, including the ones nobody piloted it on.Scoring the pilot ward. Also missing cascade: a wrong note feeds coding, research and the next clinician.
D5 Sensitivity and UncertaintyClinical data is 4. Generative summarization of clinical content carries uncertainty 3 or 4 because omission is the dangerous failure and omission is hard to detect.Averaging sensitivity and uncertainty. Take the higher.

4. One system, end to end

The overlay above is a map. This is one route across it: a single patient encounter followed from telemetry to outcome, with the record or control that attaches at each step.

Worked example — deterioration prediction on an inpatient ward D1 4 · D2 2–3 · D3 4 · D4 4 · D5 4 → IMPACT 4 · CONTROL DEFICIT 4 → TIER 4 IN THE WORLD IN THE RECORD, AND WHAT WATCHES IT Vitals and labs stream from the electronic record Lineage for every extract; an EHR upgrade that changes a field is a drift event Model scores risk and raises an alert Site authorization, not just device clearance: this population, this workflow Nurse sees the alert among the others Alert burden measured per clinician per shift; fatigue is a control failure Clinician assesses and escalates or stands down Acceptance- rate ceiling agreed at G2; breach re- scores D2 rather than being noted Outcome recorded; case reviewed Drift watch: acceptance rate, subgroup performance, population shift, vendor update
Why this is Tier 4 even though a clinician decides. D3 is at its ceiling because a missed deterioration cannot be undone, and D2 is read from measured acceptance rather than from the interface. The two controls that matter sit in columns three and four, and both are about the human — the one part of the system that the model's own metrics never measure.

5. What each gate adds

Additions only. Everything in the base gate definitions still applies; see the gate checklists for the common set.

Gate Sector addition Why it is here
G1State the intended clinical use in the words a clinician would use, and state explicitly what the system is not for. Confirm whether the software is a device under the applicable regime.Off-label drift starts with an intended-use statement written by the people funding the project.
G2Population comparison: the validation population against the deployment population, on the characteristics that matter clinically. Named data owner for every source, including the electronic record extract.A model validated elsewhere is a hypothesis at your site until this comparison exists.
G3Site-level authorization by a named clinical owner, separate from device clearance. Agreed acceptance-rate ceiling and monitoring plan. Alert burden estimate for anything that interrupts.Clearance says the device is safe in general. G3 says this deployment, on this population, with this workflow, is safe here.
G4Any model update outside the predetermined change control plan is Major. Any workflow change that alters where the clinician can intervene is an autonomy change.The workflow is part of the system, and it is changed by people who never touch the model.
G5Records retention to the clinical standard, and a preserved ability to explain a historical recommendation for the duration of the liability window.Clinical questions arrive years later, and they are about a version that no longer exists.

6. Controls and evidence worth adding

Control Where it attaches Evidence it produces
Named Clinical Safety Officer or equivalent as the accountable ownerDeployment Authorization RecordA named clinician, per site, with the authority to suspend
Acceptance-rate and override monitoring against a pre-agreed ceilingRuntime signal; assurance cycleMeasured acceptance rate by unit and shift, with the re-score trigger stated
Subgroup performance monitoring on clinically relevant strataControl Matrix; assurance cyclePerformance by stratum with the sample sizes that make it meaningful
Alert burden and fatigue measurement for interrupting systemsRuntime signalAlerts per clinician per shift, and the dismissal rate trend
Documentation accuracy sampling for generated notesControl MatrixSampled notes reviewed against source audio or encounter, with the omission rate recorded
Suspension runbook that a clinical lead can execute without engineeringKill-switch record at Tier 3–4A tested procedure, with the last test date

7. Runtime signals to wire first

Control Plane onboarding order matters more than coverage in the first year (Chapter 18). These are the signals that earn their place earliest in this sector.

Signal Drift category Suggested response
Acceptance rate rises above the agreed ceilingBehavioral, with a classification consequenceRe-score D2 and review the oversight design; the fix is usually the workflow, not the model
Deployment population diverges from the validation populationAI-model behavioralTrigger the assurance re-run at Tier 3–4; a population shift is a silent performance change
Electronic record upgrade changes an extract, a code set or a field meaningData lineage and governanceSame-day owner alert; most clinical AI breakage originates here rather than in the model
Alert volume per clinician exceeds the agreed burdenBehavioralTreat as a control failure and route to the clinical owner; fatigue disables the human control
Vendor pushes a model update to a hosted clinical serviceAI-model behavioralVerify it falls inside the cleared change plan; if not, it is a Major change and the deployment authorization is void until re-taken

8. Failure modes this sector produces

Advisory in name

The tell. The system is scored D2 = 1, the interface has an accept button, and the accept rate is 97% at 3 a.m.

The response. Instrument acceptance before deployment and agree the ceiling at G2. Where acceptance is structurally certain, design the oversight point to require something the system did not supply — a reason code, a choice between genuinely different options.

The scribe that writes the record

The tell. Ambient documentation is procured as an efficiency measure by an operations team, and the generated note becomes the clinical record other decisions rely on.

The response. Score it on the decisions its output shapes. Sample for omission specifically: fluent notes that leave something out read better than accurate ones.

Cleared once, changed continuously

The tell. A cleared model is updated by the supplier under a managed service, and nobody at the site can say which version produced last month's recommendations.

The response. Version identity is a G3 requirement and a runtime signal. If the supplier cannot tell you which version ran on a given date, that is a procurement finding, not a technical one.

Governance stops at the department that bought it

The tell. A tool piloted in one unit spreads by word of mouth into three others with different populations.

The response. Population change is a Major change. Detect spread from access telemetry, not from the project plan.

Equity monitored once, at validation

The tell. Subgroup performance was assessed in the original study and never again at the site.

The response. Subgroup monitoring belongs in the assurance cycle with a named owner, sample sizes agreed in advance, and a stated action when a stratum degrades.

9. A ninety-day start

If the sector is yours and the framework is new, this is the order that produces something defensible fastest. It assumes one part-time architect and one risk lead, not a programme.

  1. Inventory every AI system touching a patient pathway, including features inside the electronic record and anything procured by a department rather than by IT.
  2. Name a clinical owner for each. Systems without one are the first backlog, ahead of any scoring work.
  3. Score the five highest-consequence systems with the clinical safety function in the room, and record the acceptance ceiling for each.
  4. Compare validation and deployment populations for the top two. Expect this to take longer than the scoring did.
  5. Instrument acceptance rate and alert burden on the two most-used systems.
  6. Write the suspension runbook for the highest-tier system and test it once, with the clinical lead executing it.
  7. Wire one drift signal: electronic-record extract change. It breaks more clinical AI than model change does.
  8. Run a G3 dry run on a system already live, and treat the gap list as the programme plan.