K · Practical guide

Practical Guide: Banking and Capital Markets

Where an existing model risk function already does half the work, where the generic anchors read too low, and what the payments and surveillance systems need that a credit model does not.

Typical top tier
Tier 4 — credit decisioning, payment-acting agents, execution algorithms
Already-owned ground
Model risk management, operational resilience, records retention
Hardest gate
G3 — because independent validation and deployment authorization are different decisions
First record to fix
The single AI inventory, reconciled with the model inventory

1. Where the risk actually concentrates

Banking is the sector most likely to have something already. A model risk management function, a validation team, an inventory, a challenger process and a documented approval chain usually exist before any of this arrives. The overlay is therefore mostly about interfacing, not building. The framework's contribution is the systems that fall outside the model risk perimeter — generative assistants drafting customer communications, retrieval systems grounded on policy documents, and agents that act in operations — together with the runtime comparison that model validation traditionally does not perform.

Risk concentrates in four places. Decisions about access to money: credit, affordability, account opening and closure, collections. Systems that move or block money in real time: payments, fraud interdiction, execution. Systems whose failure is a reporting failure rather than a customer harm: transaction monitoring, surveillance, regulatory reporting. And customer-facing generative systems whose failure mode is a statement the firm is bound by.

The last of these is the one governance programmes consistently under-read. A retail assistant that tells a customer their arrangement fee is waived has made a statement about a product. Whether that was a hallucination is a technical fact; whether the firm honours it is a conduct question that will be answered by someone senior, at speed, in public.

Intake onboarding, KYC, document capture Tier 2 Credit affordability, limits, collections Tier 4 Payments fraud and sanctions interdiction Tier 4 Servicing assistants, correspondence Tier 3 Control surveillance, reporting, reconciliation Tier 3
Where the tiers sit across a bank's AI estate. The badges are the tier the sector's systems typically reach, not a shortcut around scoring. Note the pattern: tier peaks where money moves or is refused, and the servicing column — usually the least governed — sits only one step below it.

Indicative classification of the systems this sector keeps building. The scores are illustrative, not authoritative: they show how the anchors in Chapter 6 read against sector facts. Score your own system; do not copy a row.

Typical system Illustrative D1–D5 Tier What is usually mis-scored
Credit decisioning or affordability assessmentD1 4 · D2 2–3 · D3 3 · D4 4 · D5 3Tier 3–4 (safety-override floor applies at D1 = 4)D2 scored 2 on the strength of a reviewer who declines under 1% of proposals
Real-time fraud or sanctions interdictionD1 3–4 · D2 4 · D3 3 · D4 4 · D5 3Tier 4D3 read as reversible because a block can be lifted — the missed payment cannot
Transaction-monitoring alert triage and suppressionD1 3–4 · D2 3 · D3 3 · D4 3 · D5 3Tier 3–4D1 scored on customer harm only; the harm is a report never filed
Execution or trading strategy with a learned componentD1 4 · D2 4 · D3 4 · D4 3–4 · D5 3Tier 4Nothing — but the record often lives only in the algo governance file, not the AI inventory
Collections, arrears and forbearance prioritizationD1 3–4 · D2 2 · D3 3 · D4 3 · D5 3Tier 3Vulnerability is treated as a data attribute rather than a consequence multiplier on D1
Retail customer assistant answering product questionsD1 3 · D2 2–3 · D3 3 · D4 4 · D5 3Tier 3D3 scored 1 because the chat log can be deleted; the customer has already read it
Reconciliation or journal-posting agentD1 3 · D2 4 · D3 3 · D4 3 · D5 2Tier 4See the worked dispute in Chapter 7, Example D
KYC document extraction feeding a human reviewerD1 2–3 · D2 2 · D3 2 · D4 3 · D5 4Tier 2–3D5 under-read: identity documents are among the most sensitive inputs the firm holds
Analyst research summarization on internal documentsD1 2 · D2 1 · D3 2 · D4 2 · D5 3–4Tier 2Material non-public information in the grounding corpus, scored as ordinary confidentiality
Drafting of model documentation or validation evidenceD1 3 · D2 1 · D3 2 · D4 2 · D5 3Tier 2, with a control additionTreated as a productivity tool although its output is the firm's assurance record

2. The regulatory interface

Regulatory note. The regimes below are named so each IRGF record can be pointed at the obligation it evidences, not to restate them. Applicability, thresholds and commencement dates differ by jurisdiction and several have moved during implementation. Nothing here is legal advice: confirm the current position with your own counsel, and record the answer in the Regulatory Overlay Reference so it is checkable later.

IRGF does not restate any of these obligations. It gives each one a record that carries the evidence, an owner, and a trigger that reopens it when the obligation or the system changes.

Regime or standard What it obliges in practice IRGF record that carries the evidence
Model risk management supervision (for example SR 11-7 / OCC 2011-12 in the US, PRA SS1/23 in the UK)A model inventory, tiering, independent validation proportionate to tier, ongoing monitoring and a documented owner. Most AI systems in scope of IRGF are also models under these definitions.Risk Classification Record reconciled with the model inventory; validation output referenced from the AI Assurance Summary rather than repeated
Automated decision-making and profiling rules (for example GDPR Article 22 and equivalents)A meaningful human oversight point, an explanation of the logic involved, and a route to contest a decision.Human oversight point confirmed at G3 in the Deployment Authorization Record; D2 anchor evidence from override-rate instrumentation
Adverse action and fair lending obligations (for example ECOA/Regulation B, FCRA, EBA loan origination guidelines)Specific, accurate reasons for a decline; no reliance on a reason the firm cannot substantiate; monitoring of outcomes.Reason-code generation recorded as a control in the Control Matrix, mapped to the D5 uncertainty score
EU AI Act, where it appliesCreditworthiness assessment of natural persons sits in the Annex III high-risk list, with obligations on risk management, data governance, logging, human oversight, accuracy and post-market monitoring. Life and health insurance pricing is treated similarly.Regulatory Overlay Reference mapping each affected system; logging obligations satisfied from the audit link of the agent chain
Digital operational resilience rules (for example DORA in the EU)ICT third-party register, concentration risk assessment, exit strategies, incident classification and reporting, resilience testing.Vendor fields of the AI-extended ADR; Drift/Alert Record routing into the existing incident taxonomy
Algorithmic trading requirements (for example MiFID II RTS 6)Pre-deployment testing in a segregated environment, kill functionality, annual self-assessment, documented change control.Kill-switch conditions pre-approved at G3; annual self-assessment aligned to the G4 re-authorization cycle rather than run as a separate exercise
Risk data aggregation principles (BCBS 239)Accuracy, completeness and traceability of the data feeding risk reporting, including lineage.Data Lineage and Sensitivity Record, which is the same artifact the grounding sources need anyway
Conduct and fair-treatment regimes (for example FCA Consumer Duty)Evidence that communications support understanding, that outcomes are monitored, and that foreseeable harm is avoided — including for customers in vulnerable circumstances.Outcome monitoring registered as a runtime signal; vulnerability handling recorded as a D1 justification rather than a data field

3. Calibrating the five dimensions

The dimensions do not change. What changes is what a 3 and a 4 look like when the subject matter is this sector, and which reading an assessor under delivery pressure reaches for first.

Dimension How to read it here The mis-score to watch for
D1 Decision ConsequenceAny output that determines access to credit, to an account, or to a customer's own funds starts at 3. Denial, withdrawal or closure, and anything that produces an adverse record a third party will rely on, is 4.Scoring on the firm's loss rather than the customer's. A decision that costs the bank nothing and the customer their tenancy application is a 4.
D2 AutonomyOperational review under volume targets is the sector's characteristic autonomy trap. A credit officer processing 300 proposals a day is not exercising judgment on each.Recording 2 for “human approves each” without measuring the override rate. Instrument it before the score is countersigned, not after.
D3 Reversibility DeficitMoney that has moved, a report that was not filed in its window, a bureau record, and a statement made to a customer are all effectively irreversible regardless of what the ledger can undo.Mechanism reversibility. A reversal entry restores the balance; it does not un-file the regulatory return that used the wrong number.
D4 Exposure and ScaleCount the reach of the records the system writes into, not the number of people who log in. A system feeding the general ledger or a regulatory return has enterprise reach with four users.Scoring 2 for “internal, finance team only” when the output lands in the statutory accounts.
D5 Sensitivity and UncertaintyFinancial transaction data, identity documents and material non-public information each sit at 3 or above. Take the higher of the sensitivity and uncertainty sub-scores; do not average them.Treating MNPI as ordinary confidential data because the classification scheme has no category for it.

4. One system, end to end

The overlay above is a map. This is one route across it: a single system, followed from the moment data arrives to the moment the consequence leaves the firm, with the record that attaches at each step.

Worked example — affordability assessment on a personal loan application D1 4 · D2 2 · D3 3 · D4 4 · D5 3 → IMPACT 4 · CONTROL DEFICIT 3 → TIER 4 IN THE WORLD IN THE RECORD, AND WHAT WATCHES IT Application and bureau data arrive Canvas names the decision and the bureau feeds; provisional D1–D5 recorded at G1 Model scores affordability and proposes a decline Classification countersigned; D1 = 4 because a decline denies access to credit Credit officer reviews the proposed decline Oversight design note: what the officer sees, and the override-rate floor agreed at G2 Decision and reasons issued to the customer Reason-code control in the Control Matrix; authorization names the accountable owner Outcome reported to the bureau Drift watch: bureau feed change, override rate, supplier model version
One system, from application to bureau record. The tier is set by the fourth column, not the first: the moment a wrong output denies someone credit and is reported onward, D1 and D3 are both at their ceiling. Note where the governance work actually lands — two of the five columns are about the human who reviews, because that is the control most likely to be nominal.

5. What each gate adds

Additions only. Everything in the base gate definitions still applies; see the gate checklists for the common set.

Gate Sector addition Why it is here
G1State whether the system is a model under the firm's existing model risk policy, and record the answer. State whether the grounding corpus can contain material non-public information.Both answers determine which existing process owns the work. Deciding late means either duplicate review or an uncovered system.
G2Lineage to BCBS 239 standard for anything feeding risk or regulatory reporting. Vendor concentration assessment where the model, the hosting and the retrieval layer come from one supplier.Concentration is invisible at system level and only appears when the portfolio is looked at once.
G3Independent validation sign-off referenced, not repeated. Reason-code coverage evidenced for any declining decision. Kill-switch conditions for anything acting on payments or orders.Validation answers “is the model sound”. G3 answers “is this safe to run here, now, at this tier”. Collapsing them loses the second question.
G4Any vendor model version change at Tier 3–4 is Material by default. Challenger promotion is a Major change, not a tuning exercise.The most common uncontrolled change in banking AI is a supplier upgrading a model under a managed service.
G5Retention set by the longest applicable record-keeping obligation, and evidence retrievable in a form a regulator can read after the platform is gone.Retirement usually removes the ability to answer questions about decisions the system made while it ran.

6. Controls and evidence worth adding

Control Where it attaches Evidence it produces
Single AI and model inventory, reconciled monthlyRisk Classification Record, model inventoryA reconciliation report showing zero systems in one inventory and not the other
Override-rate instrumentation on every D2 = 1 or 2 system with D1 ≥ 3Runtime telemetry, reviewed at the first assurance cycleA measured override rate against a threshold agreed in advance
Reason-code generation and substantiation checkControl Matrix, mapped to the adverse action obligationSampled decisions with the reason given and the factor that actually drove it
Value, counterparty and rate limits for any acting agentAgent Card, enforced in the policy engineMachine-readable boundary, comparable against the IAM grant
Segregation of material non-public information from retrieval corporaData Lineage and Sensitivity Record; retrieval index buildSource-level classification and an index build log that excludes restricted classes
Dual control on any change to an agent's authority boundaryChange record, policy repositoryTwo named approvers and a diff of the boundary
Alert suppression rate monitoring for transaction monitoringRuntime signal into the Control PlaneSuppression rate by typology, with a change threshold that triggers re-score

7. Runtime signals to wire first

Control Plane onboarding order matters more than coverage in the first year (Chapter 18). These are the signals that earn their place earliest in this sector.

Signal Drift category Suggested response
Override rate on a reviewed decision system falls below the agreed floorBehavioral, with a classification consequenceRe-score D2 rather than note it; the recorded tier is wrong, not the behaviour
Supplier changes the underlying model version of a managed serviceAI-model behavioralTier 3–4: treat as Material change and route to G4. Tier 1–2: log and batch-review
Retrieval corpus gains a source above its approved sensitivityData lineage and governanceBlock the index build if the pipeline can enforce it; otherwise same-day owner alert and re-score of D5
Agent posting outside its approved account range or above its value limitAgent boundary divergenceIncident at any tier: suspend, revoke, preserve state, notify the Agent Owner and AI Governance Lead
Alert suppression rate moves more than the agreed band, in either directionBehavioralRoute to the financial crime team with the model owner, not to the architect
Consumer or conduct rules change in a mapped jurisdictionPolicy and regulatoryRe-open compliance assurance for every system in the overlay mapping, without waiting for a release

8. Failure modes this sector produces

Two inventories, two review queues, one system

The tell. Model risk asks for a validation pack and AI governance asks for an assurance summary, and the delivery team writes the same content twice with different words. Within two quarters the two records disagree and nobody knows which is authoritative.

The response. Decide the interface once: the model inventory is the register, the Risk Classification Record is the tiering, and the assurance summary references the validation report rather than restating it. Reconcile monthly and publish the exceptions.

The assistant that quietly became a decision system

The tell. A generative assistant deployed to “help advisers draft responses” is, six months later, producing the response that goes out with a one-click send. Nothing was reclassified because nothing was rebuilt.

The response. Autonomy change is a change type in its own right (Chapter 17). Detect it from the interaction telemetry — the ratio of edits to sends — not from the change log.

Validation evidence written by the system under validation

The tell. An LLM drafts the model documentation, the validation summary and the challenge log. The reviewer reads fluent text that agrees with itself and signs.

The response. Permit drafting; prohibit the drafting system from being the source of the evidence. Require the challenge log to record at least one substantive disagreement and its resolution, authored by a named person.

Tier laundering through the human-in-the-loop claim

The tell. Every high-consequence system arrives at G2 with D2 = 2 and a screenshot of an approval button.

The response. Countersignature is the control (Chapter 7, §7.6), and the override-rate threshold must be agreed at G2, before anyone knows the answer.

Concentration discovered during an outage

The tell. Model, embeddings, vector store and the assistant framework all come from one supplier, and each system cleared G2 on its own merits.

The response. Concentration is a portfolio property. Add a standing quarterly question to the Architecture Review Board: which single supplier failure would take out the most Tier 3–4 systems, and what is the exit?

9. A ninety-day start

If the sector is yours and the framework is new, this is the order that produces something defensible fastest. It assumes one part-time architect and one risk lead, not a programme.

  1. List every AI system touching credit, payments, financial crime and customer communications. Include the ones procured as features of existing platforms — that is where the omissions are.
  2. Reconcile that list against the model inventory. Publish the two exception lists: models that are not in the AI inventory, and AI systems nobody has tiered.
  3. Agree the interface with model risk in one page: who tiers, who validates, which record is authoritative, and what G3 adds that validation does not.
  4. Score the five highest-consequence systems with an independent countersigner present, using the classification checklist.
  5. Instrument override rate on the two reviewed decision systems with the highest volume. Agree the floor before you see the number.
  6. Write authority boundaries for every agent that touches a ledger, a payment rail or an order, using the Agent Card. Boundaries first, kill-switches second.
  7. Wire two drift signals only: grounding-source change on the highest-tier retrieval system, and supplier model version change across the estate.
  8. Run one G3 dry run on a system already in production. The gap list it produces is your real backlog.

The order matters. Firms that begin with tooling spend the first year integrating a repository and the second discovering that the tiers in it were never defensible.