Practical Guide: Government and Public Sector
Administrative law already requires most of what the framework asks for. The difference is that here the evidence has to survive publication, appeal and a change of minister.
- Typical top tier
- Tier 4 — benefits, enforcement, immigration, child and adult safeguarding
- Already-owned ground
- Duty to give reasons, appeal routes, records law, equality duties
- Hardest gate
- G1 — because the decision to use AI at all is itself reviewable
- First record to fix
- A published, plain-language description of every citizen-facing system
1. Where the risk actually concentrates
Public bodies operate under a constraint no commercial organization shares: the reasoning behind a decision may have to be explained to the person affected, defended before a tribunal, disclosed under access-to-information law, and published in a transparency register. Governance evidence that is adequate internally may be inadequate the moment it becomes public, and it will become public.
That changes the artifact set's centre of gravity. The record that matters most is not the assurance summary but the explanation: what the system does, what data it uses, what a person can do if they disagree, and who is accountable. IRGF's derived AI System Card is the natural vehicle, and in this sector it should be written for a citizen rather than for an architect.
The second difference is that eligibility, enforcement and safeguarding decisions concentrate harm on people with the least capacity to contest them. A private-sector error costs money; a public-sector error can cost someone their income, their liberty or their child's placement. D1 anchors read accordingly, and the framework's safety-override floor will apply more often here than anywhere else in this guide.
Indicative classification of the systems this sector keeps building. The scores are illustrative, not authoritative: they show how the anchors in Chapter 6 read against sector facts. Score your own system; do not copy a row.
| Typical system | Illustrative D1–D5 | Tier | What is usually mis-scored |
|---|---|---|---|
| Benefit or entitlement eligibility determination | D1 4 · D2 2–4 · D3 4 · D4 4 · D5 4 | Tier 4 | D3 read as reversible because an appeal exists; the arrears arrive months after the eviction |
| Fraud and error risk scoring on claims | D1 4 · D2 3 · D3 4 · D4 4 · D5 4 | Tier 4 | A referral is scored as advisory although it triggers an investigation with real consequences |
| Immigration, visa or border risk triage | D1 4 · D2 3 · D3 4 · D4 4 · D5 4 | Tier 4 | Prioritization is treated as workflow; it determines who is examined and who is not |
| Child or adult safeguarding risk prediction | D1 4 · D2 2 · D3 4 · D4 3 · D5 4 | Tier 4 | Advisory scoring; in practice the score frames every subsequent professional judgment |
| Enforcement or inspection targeting | D1 3–4 · D2 3 · D3 3 · D4 4 · D5 3 | Tier 3–4 | Selection effects compound: the data records what was inspected, not what was wrong |
| Public-facing information assistant on rules and entitlements | D1 3–4 · D2 3 · D3 4 · D4 4 · D5 3 | Tier 3–4 | Wrong guidance about a deadline is not recoverable by correcting the page |
| Case-worker drafting and correspondence generation | D1 3 · D2 2 · D3 3 · D4 4 · D5 4 | Tier 3 | The letter is the decision as far as the recipient is concerned |
| Procurement, HR or internal document processing | D1 2–3 · D2 2 · D3 2 · D4 3 · D5 3 | Tier 2 | Recruitment screening is not internal: it is an Annex III-style employment decision |
| Traffic, planning or scheduling optimization | D1 2–3 · D2 3 · D3 2 · D4 4 · D5 2 | Tier 2–3 | Distributional effects across communities are the risk, not average performance |
2. The regulatory interface
Regulatory note. The regimes below are named so each IRGF record can be pointed at the obligation it evidences, not to restate them. Applicability, thresholds and commencement dates differ by jurisdiction and several have moved during implementation. Nothing here is legal advice: confirm the current position with your own counsel, and record the answer in the Regulatory Overlay Reference so it is checkable later.
IRGF does not restate any of these obligations. It gives each one a record that carries the evidence, an owner, and a trigger that reopens it when the obligation or the system changes.
| Regime or standard | What it obliges in practice | IRGF record that carries the evidence |
|---|---|---|
| Administrative law duties in the relevant jurisdiction | A lawful basis for the decision, reasons the affected person can understand, a route to challenge, and no unlawful fetter on discretion by an automated process. | Reasons capability recorded as a control; oversight point at G3; discretion preserved in the workflow design, and evidenced |
| Central government AI policy (for example the OMB memoranda governing US federal agency use and acquisition of AI, which define a high-impact or rights-impacting category with minimum practices, inventories and a senior accountable official) | Use-case inventory, impact assessment, testing, human oversight, and public reporting for the highest-impact uses. | The AI inventory is the Risk Classification Record set; the senior accountable official is the escalation point in the RACI |
| Algorithmic transparency standards (for example the UK Algorithmic Transparency Recording Standard for central government) | A published record describing the tool, its purpose, data, oversight and contact route. | Published from the derived AI System Card, compiled by reference rather than authored separately |
| EU AI Act, where it applies | Annex III covers essential public services, law enforcement, migration and justice uses; public bodies deploying high-risk systems face a fundamental rights impact assessment duty alongside the provider obligations. | Impact assessment attached to the Use-Case Canvas at S1 and refreshed at every Major change |
| Data protection law for public authorities | Lawful basis, purpose limitation, data minimisation, and rules on automated decisions with legal or similarly significant effect. | Lineage record; oversight point; D2 anchored against the legal-effect test rather than the interface |
| Equality and non-discrimination duties | Assessment of effects on protected groups before adoption, and monitoring after. | Equality assessment referenced by the Control Matrix; subgroup monitoring as a runtime signal |
| Records and access-to-information law | Retention schedules, and disclosure of information including, in some regimes, the logic of automated processing. | Retention set at G5; evidence written on the assumption it will be read outside the organization |
3. Calibrating the five dimensions
The dimensions do not change. What changes is what a 3 and a 4 look like when the subject matter is this sector, and which reading an assessor under delivery pressure reaches for first.
| Dimension | How to read it here | The mis-score to watch for |
|---|---|---|
| D1 Decision Consequence | Anything affecting income, liberty, immigration status, housing, care or a child's placement is 4. There is no proportionality argument that makes it 3. | Scoring the administrative cost of an error rather than its effect on the person. |
| D2 Autonomy | A score presented to a caseworker with a default action is not advisory. Where the recommendation is followed in the overwhelming majority of cases, effective autonomy is 3. | The caseworker-in-the-loop claim, used to keep systems out of the high-impact category. |
| D3 Reversibility Deficit | Appeal is not reversal. Score against elapsed harm: a benefit restored after four months did not undo the four months. | Treating a statutory appeal route as making the consequence reversible. |
| D4 Exposure and Scale | Population-wide by default. Add a cascade point where one determination feeds another agency's decision. | Scoring the pilot region. Also missing inter-agency data flows, which are the sector's real scale factor. |
| D5 Sensitivity and Uncertainty | Benefits, health, immigration, criminal justice and children's data are 4. Generative explanation of a decision carries uncertainty 3 because a fluent wrong reason is worse than no reason. | Treating administrative data as low sensitivity because it is not health data. |
4. One system, end to end
The overlay above is a map. This is one route across it: a single claim followed from submission to appeal, with the record or control that attaches at each step.
5. What each gate adds
Additions only. Everything in the base gate definitions still applies; see the gate checklists for the common set.
| Gate | Sector addition | Why it is here |
|---|---|---|
| G1 | Record why an automated approach is appropriate for a decision of this kind, and what the non-AI alternative would cost in fairness terms, not only in money. Confirm the legal power to process this data for this purpose. | In this sector the decision to automate is itself challengeable, and the reasoning has to exist from the start rather than be reconstructed. |
| G2 | Data provenance including how historical decisions in the training data were made, and whether they encode past practice you would not defend today. | Public sector training data is a record of prior administrative behaviour, including its errors. |
| G3 | A citizen-readable description ready to publish, a reasons capability that produces the actual operative factors, and a named accountable official who is not the project sponsor. | Publication is not a communications task added afterwards; it is evidence that the authorization was sound. |
| G4 | Any threshold change is Material where it changes who is selected, referred or refused. Publication record updated as part of the change, not after it. | Thresholds are adjusted operationally under performance pressure, and they are exactly the parameter a tribunal will ask about. |
| G5 | Retention under the records schedule, and preserved ability to explain a decision made years earlier to a tribunal or ombudsman. | Challenges arrive long after the system is switched off. |
6. Controls and evidence worth adding
| Control | Where it attaches | Evidence it produces |
|---|---|---|
| Published plain-language system description, kept current | Derived AI System Card; transparency register | A public record with a date and a contact route |
| Operative-reasons capability, tested against real cases | Control Matrix | Sampled decisions where the stated reason matches the factor that actually drove the outcome |
| Human decision-maker with genuine discretion, evidenced | Deployment Authorization Record; workflow design | Measured departure rate from the recommendation, and cases where discretion was exercised |
| Equality and distributional monitoring by group and region | Assurance cycle | Outcome rates by group with pre-agreed action thresholds |
| Appeal and complaint feedback loop into assurance | Runtime signal | Overturn rate on appeal, treated as a performance measure of the system |
| Selection-effect review for targeting systems | Assurance cycle | Analysis of whether the system is learning from its own prior selections |
7. Runtime signals to wire first
Control Plane onboarding order matters more than coverage in the first year (Chapter 18). These are the signals that earn their place earliest in this sector.
| Signal | Drift category | Suggested response |
|---|---|---|
| Overturn rate on appeal rises | Behavioral | Treat as a primary quality signal, not a legal statistic; route to the system owner and the policy owner |
| Caseworker departure rate from recommendations falls toward zero | Behavioral, with a classification consequence | Re-score D2; the oversight control has stopped operating |
| Eligibility rules or policy change | Policy and regulatory | Immediate compliance review: rule change without model change is the sector's most common silent non-compliance |
| Selection or referral rate changes by group or region | Behavioral | Escalate on the pre-agreed threshold, with the policy owner present |
| Inter-agency data feed changes definition or coverage | Data lineage and governance | Same-day owner alert; a changed field meaning silently changes eligibility |
8. Failure modes this sector produces
Transparency written after authorization
The tell. The published description is drafted by a communications team from a technical document, months after go-live, and it does not match what the system does.
The response. Draft the public description at G3 as evidence. If it cannot be written honestly and simply, the authorization is not ready.
The recommendation that is really a decision
The tell. A caseworker may depart from the score, but departure requires a written justification and a manager's approval, while acceptance requires a click.
The response. Symmetric friction. If departing costs more than accepting, the score is the decision and the system should be tiered as such.
Learning from the enforcement record
The tell. A targeting model trained on past inspections learns where inspectors used to go, and the estate confirms its own history.
The response. Add a selection-effect review to the assurance cycle and reserve a random sample of unselected cases as a control group.
Pilot with consent, deploy without
The tell. A voluntary pilot with an opt-out becomes a mandatory pathway once the efficiency case is made.
The response. The change from voluntary to mandatory is a Major change to D2 and D4 and re-enters authorization.
Reasons that are fluent and wrong
The tell. A generative explanation layer produces plausible reasons that are not the operative factors.
The response. Generate reasons from the decision logic, not from a language model reading the outcome. If that is not possible, the system cannot support a decision requiring reasons.
9. A ninety-day start
If the sector is yours and the framework is new, this is the order that produces something defensible fastest. It assumes one part-time architect and one risk lead, not a programme.
- Inventory every system that affects a citizen-facing decision, including scoring, prioritization and correspondence tools. Departmental spreadsheets with models in them count.
- Publish a plain-language description of the three highest-impact ones, even if the register does not yet require it. It surfaces gaps faster than any internal review.
- Score those three with legal and policy owners present, and expect the safety-override floor to apply.
- Measure caseworker departure rates on any system claiming human decision-making.
- Check the reasons capability on twenty real cases: does the stated reason match the operative factor?
- Map eligibility and policy rules to the systems that implement them, so a rule change has a known blast radius.
- Set up overturn-on-appeal reporting into the assurance cycle.
- Run a tribunal-style dry run: reconstruct one decision from six months ago end to end.