K · Practical guide

Practical Guide: Government and Public Sector

Administrative law already requires most of what the framework asks for. The difference is that here the evidence has to survive publication, appeal and a change of minister.

Typical top tier
Tier 4 — benefits, enforcement, immigration, child and adult safeguarding
Already-owned ground
Duty to give reasons, appeal routes, records law, equality duties
Hardest gate
G1 — because the decision to use AI at all is itself reviewable
First record to fix
A published, plain-language description of every citizen-facing system

1. Where the risk actually concentrates

Public bodies operate under a constraint no commercial organization shares: the reasoning behind a decision may have to be explained to the person affected, defended before a tribunal, disclosed under access-to-information law, and published in a transparency register. Governance evidence that is adequate internally may be inadequate the moment it becomes public, and it will become public.

That changes the artifact set's centre of gravity. The record that matters most is not the assurance summary but the explanation: what the system does, what data it uses, what a person can do if they disagree, and who is accountable. IRGF's derived AI System Card is the natural vehicle, and in this sector it should be written for a citizen rather than for an architect.

The second difference is that eligibility, enforcement and safeguarding decisions concentrate harm on people with the least capacity to contest them. A private-sector error costs money; a public-sector error can cost someone their income, their liberty or their child's placement. D1 anchors read accordingly, and the framework's safety-override floor will apply more often here than anywhere else in this guide.

Policy modelling, consultation analysis Tier 2 Eligibility benefits, licences, entitlement Tier 4 Enforcement fraud, targeting, inspection Tier 4 Casework triage, correspondence, drafting Tier 3 Corporate recruitment, procurement Tier 2-3
Harm concentrates where contestability is weakest. The two Tier 4 columns decide access to income, status and liberty for people with the least capacity to challenge a wrong answer, which is why the safety-override floor applies more often in this sector than in any other in this guide.

Indicative classification of the systems this sector keeps building. The scores are illustrative, not authoritative: they show how the anchors in Chapter 6 read against sector facts. Score your own system; do not copy a row.

Typical system Illustrative D1–D5 Tier What is usually mis-scored
Benefit or entitlement eligibility determinationD1 4 · D2 2–4 · D3 4 · D4 4 · D5 4Tier 4D3 read as reversible because an appeal exists; the arrears arrive months after the eviction
Fraud and error risk scoring on claimsD1 4 · D2 3 · D3 4 · D4 4 · D5 4Tier 4A referral is scored as advisory although it triggers an investigation with real consequences
Immigration, visa or border risk triageD1 4 · D2 3 · D3 4 · D4 4 · D5 4Tier 4Prioritization is treated as workflow; it determines who is examined and who is not
Child or adult safeguarding risk predictionD1 4 · D2 2 · D3 4 · D4 3 · D5 4Tier 4Advisory scoring; in practice the score frames every subsequent professional judgment
Enforcement or inspection targetingD1 3–4 · D2 3 · D3 3 · D4 4 · D5 3Tier 3–4Selection effects compound: the data records what was inspected, not what was wrong
Public-facing information assistant on rules and entitlementsD1 3–4 · D2 3 · D3 4 · D4 4 · D5 3Tier 3–4Wrong guidance about a deadline is not recoverable by correcting the page
Case-worker drafting and correspondence generationD1 3 · D2 2 · D3 3 · D4 4 · D5 4Tier 3The letter is the decision as far as the recipient is concerned
Procurement, HR or internal document processingD1 2–3 · D2 2 · D3 2 · D4 3 · D5 3Tier 2Recruitment screening is not internal: it is an Annex III-style employment decision
Traffic, planning or scheduling optimizationD1 2–3 · D2 3 · D3 2 · D4 4 · D5 2Tier 2–3Distributional effects across communities are the risk, not average performance

2. The regulatory interface

Regulatory note. The regimes below are named so each IRGF record can be pointed at the obligation it evidences, not to restate them. Applicability, thresholds and commencement dates differ by jurisdiction and several have moved during implementation. Nothing here is legal advice: confirm the current position with your own counsel, and record the answer in the Regulatory Overlay Reference so it is checkable later.

IRGF does not restate any of these obligations. It gives each one a record that carries the evidence, an owner, and a trigger that reopens it when the obligation or the system changes.

Regime or standard What it obliges in practice IRGF record that carries the evidence
Administrative law duties in the relevant jurisdictionA lawful basis for the decision, reasons the affected person can understand, a route to challenge, and no unlawful fetter on discretion by an automated process.Reasons capability recorded as a control; oversight point at G3; discretion preserved in the workflow design, and evidenced
Central government AI policy (for example the OMB memoranda governing US federal agency use and acquisition of AI, which define a high-impact or rights-impacting category with minimum practices, inventories and a senior accountable official)Use-case inventory, impact assessment, testing, human oversight, and public reporting for the highest-impact uses.The AI inventory is the Risk Classification Record set; the senior accountable official is the escalation point in the RACI
Algorithmic transparency standards (for example the UK Algorithmic Transparency Recording Standard for central government)A published record describing the tool, its purpose, data, oversight and contact route.Published from the derived AI System Card, compiled by reference rather than authored separately
EU AI Act, where it appliesAnnex III covers essential public services, law enforcement, migration and justice uses; public bodies deploying high-risk systems face a fundamental rights impact assessment duty alongside the provider obligations.Impact assessment attached to the Use-Case Canvas at S1 and refreshed at every Major change
Data protection law for public authoritiesLawful basis, purpose limitation, data minimisation, and rules on automated decisions with legal or similarly significant effect.Lineage record; oversight point; D2 anchored against the legal-effect test rather than the interface
Equality and non-discrimination dutiesAssessment of effects on protected groups before adoption, and monitoring after.Equality assessment referenced by the Control Matrix; subgroup monitoring as a runtime signal
Records and access-to-information lawRetention schedules, and disclosure of information including, in some regimes, the logic of automated processing.Retention set at G5; evidence written on the assumption it will be read outside the organization

3. Calibrating the five dimensions

The dimensions do not change. What changes is what a 3 and a 4 look like when the subject matter is this sector, and which reading an assessor under delivery pressure reaches for first.

Dimension How to read it here The mis-score to watch for
D1 Decision ConsequenceAnything affecting income, liberty, immigration status, housing, care or a child's placement is 4. There is no proportionality argument that makes it 3.Scoring the administrative cost of an error rather than its effect on the person.
D2 AutonomyA score presented to a caseworker with a default action is not advisory. Where the recommendation is followed in the overwhelming majority of cases, effective autonomy is 3.The caseworker-in-the-loop claim, used to keep systems out of the high-impact category.
D3 Reversibility DeficitAppeal is not reversal. Score against elapsed harm: a benefit restored after four months did not undo the four months.Treating a statutory appeal route as making the consequence reversible.
D4 Exposure and ScalePopulation-wide by default. Add a cascade point where one determination feeds another agency's decision.Scoring the pilot region. Also missing inter-agency data flows, which are the sector's real scale factor.
D5 Sensitivity and UncertaintyBenefits, health, immigration, criminal justice and children's data are 4. Generative explanation of a decision carries uncertainty 3 because a fluent wrong reason is worse than no reason.Treating administrative data as low sensitivity because it is not health data.

4. One system, end to end

The overlay above is a map. This is one route across it: a single claim followed from submission to appeal, with the record or control that attaches at each step.

Worked example — eligibility determination for an income-related benefit D1 4 · D2 2 · D3 4 · D4 4 · D5 4 → IMPACT 4 · CONTROL DEFICIT 3 → TIER 4 IN THE WORLD IN THE RECORD, AND WHAT WATCHES IT Claim and cross-agency data assembled Legal power to process for this purpose recorded at G1, before design Model scores eligibility and fraud risk Training data assessed for past practice the department would not defend today Caseworker decides, or accepts the score Departure rate measured; symmetric friction so departing costs no more than accepting Decision letter issued with reasons Reasons generated from the decision logic, never from a model reading the outcome Appeal or ombudsman review Drift watch: overturn rate on appeal, rule change, selection rate by group and region
Contestability is the control. Column four is where administrative law actually bites: the reason given must be the operative factor, and a fluent wrong reason is worse than none. Column five turns the appeal system into a quality signal instead of a statistic, which is the cheapest honest measure of the model this sector has.

5. What each gate adds

Additions only. Everything in the base gate definitions still applies; see the gate checklists for the common set.

Gate Sector addition Why it is here
G1Record why an automated approach is appropriate for a decision of this kind, and what the non-AI alternative would cost in fairness terms, not only in money. Confirm the legal power to process this data for this purpose.In this sector the decision to automate is itself challengeable, and the reasoning has to exist from the start rather than be reconstructed.
G2Data provenance including how historical decisions in the training data were made, and whether they encode past practice you would not defend today.Public sector training data is a record of prior administrative behaviour, including its errors.
G3A citizen-readable description ready to publish, a reasons capability that produces the actual operative factors, and a named accountable official who is not the project sponsor.Publication is not a communications task added afterwards; it is evidence that the authorization was sound.
G4Any threshold change is Material where it changes who is selected, referred or refused. Publication record updated as part of the change, not after it.Thresholds are adjusted operationally under performance pressure, and they are exactly the parameter a tribunal will ask about.
G5Retention under the records schedule, and preserved ability to explain a decision made years earlier to a tribunal or ombudsman.Challenges arrive long after the system is switched off.

6. Controls and evidence worth adding

Control Where it attaches Evidence it produces
Published plain-language system description, kept currentDerived AI System Card; transparency registerA public record with a date and a contact route
Operative-reasons capability, tested against real casesControl MatrixSampled decisions where the stated reason matches the factor that actually drove the outcome
Human decision-maker with genuine discretion, evidencedDeployment Authorization Record; workflow designMeasured departure rate from the recommendation, and cases where discretion was exercised
Equality and distributional monitoring by group and regionAssurance cycleOutcome rates by group with pre-agreed action thresholds
Appeal and complaint feedback loop into assuranceRuntime signalOverturn rate on appeal, treated as a performance measure of the system
Selection-effect review for targeting systemsAssurance cycleAnalysis of whether the system is learning from its own prior selections

7. Runtime signals to wire first

Control Plane onboarding order matters more than coverage in the first year (Chapter 18). These are the signals that earn their place earliest in this sector.

Signal Drift category Suggested response
Overturn rate on appeal risesBehavioralTreat as a primary quality signal, not a legal statistic; route to the system owner and the policy owner
Caseworker departure rate from recommendations falls toward zeroBehavioral, with a classification consequenceRe-score D2; the oversight control has stopped operating
Eligibility rules or policy changePolicy and regulatoryImmediate compliance review: rule change without model change is the sector's most common silent non-compliance
Selection or referral rate changes by group or regionBehavioralEscalate on the pre-agreed threshold, with the policy owner present
Inter-agency data feed changes definition or coverageData lineage and governanceSame-day owner alert; a changed field meaning silently changes eligibility

8. Failure modes this sector produces

Transparency written after authorization

The tell. The published description is drafted by a communications team from a technical document, months after go-live, and it does not match what the system does.

The response. Draft the public description at G3 as evidence. If it cannot be written honestly and simply, the authorization is not ready.

The recommendation that is really a decision

The tell. A caseworker may depart from the score, but departure requires a written justification and a manager's approval, while acceptance requires a click.

The response. Symmetric friction. If departing costs more than accepting, the score is the decision and the system should be tiered as such.

Learning from the enforcement record

The tell. A targeting model trained on past inspections learns where inspectors used to go, and the estate confirms its own history.

The response. Add a selection-effect review to the assurance cycle and reserve a random sample of unselected cases as a control group.

Pilot with consent, deploy without

The tell. A voluntary pilot with an opt-out becomes a mandatory pathway once the efficiency case is made.

The response. The change from voluntary to mandatory is a Major change to D2 and D4 and re-enters authorization.

Reasons that are fluent and wrong

The tell. A generative explanation layer produces plausible reasons that are not the operative factors.

The response. Generate reasons from the decision logic, not from a language model reading the outcome. If that is not possible, the system cannot support a decision requiring reasons.

9. A ninety-day start

If the sector is yours and the framework is new, this is the order that produces something defensible fastest. It assumes one part-time architect and one risk lead, not a programme.

  1. Inventory every system that affects a citizen-facing decision, including scoring, prioritization and correspondence tools. Departmental spreadsheets with models in them count.
  2. Publish a plain-language description of the three highest-impact ones, even if the register does not yet require it. It surfaces gaps faster than any internal review.
  3. Score those three with legal and policy owners present, and expect the safety-override floor to apply.
  4. Measure caseworker departure rates on any system claiming human decision-making.
  5. Check the reasons capability on twenty real cases: does the stated reason match the operative factor?
  6. Map eligibility and policy rules to the systems that implement them, so a rule change has a known blast radius.
  7. Set up overturn-on-appeal reporting into the assurance cycle.
  8. Run a tribunal-style dry run: reconstruct one decision from six months ago end to end.