Practical Guide: Banking and Capital Markets
Where an existing model risk function already does half the work, where the generic anchors read too low, and what the payments and surveillance systems need that a credit model does not.
- Typical top tier
- Tier 4 — credit decisioning, payment-acting agents, execution algorithms
- Already-owned ground
- Model risk management, operational resilience, records retention
- Hardest gate
- G3 — because independent validation and deployment authorization are different decisions
- First record to fix
- The single AI inventory, reconciled with the model inventory
1. Where the risk actually concentrates
Banking is the sector most likely to have something already. A model risk management function, a validation team, an inventory, a challenger process and a documented approval chain usually exist before any of this arrives. The overlay is therefore mostly about interfacing, not building. The framework's contribution is the systems that fall outside the model risk perimeter — generative assistants drafting customer communications, retrieval systems grounded on policy documents, and agents that act in operations — together with the runtime comparison that model validation traditionally does not perform.
Risk concentrates in four places. Decisions about access to money: credit, affordability, account opening and closure, collections. Systems that move or block money in real time: payments, fraud interdiction, execution. Systems whose failure is a reporting failure rather than a customer harm: transaction monitoring, surveillance, regulatory reporting. And customer-facing generative systems whose failure mode is a statement the firm is bound by.
The last of these is the one governance programmes consistently under-read. A retail assistant that tells a customer their arrangement fee is waived has made a statement about a product. Whether that was a hallucination is a technical fact; whether the firm honours it is a conduct question that will be answered by someone senior, at speed, in public.
Indicative classification of the systems this sector keeps building. The scores are illustrative, not authoritative: they show how the anchors in Chapter 6 read against sector facts. Score your own system; do not copy a row.
| Typical system | Illustrative D1–D5 | Tier | What is usually mis-scored |
|---|---|---|---|
| Credit decisioning or affordability assessment | D1 4 · D2 2–3 · D3 3 · D4 4 · D5 3 | Tier 3–4 (safety-override floor applies at D1 = 4) | D2 scored 2 on the strength of a reviewer who declines under 1% of proposals |
| Real-time fraud or sanctions interdiction | D1 3–4 · D2 4 · D3 3 · D4 4 · D5 3 | Tier 4 | D3 read as reversible because a block can be lifted — the missed payment cannot |
| Transaction-monitoring alert triage and suppression | D1 3–4 · D2 3 · D3 3 · D4 3 · D5 3 | Tier 3–4 | D1 scored on customer harm only; the harm is a report never filed |
| Execution or trading strategy with a learned component | D1 4 · D2 4 · D3 4 · D4 3–4 · D5 3 | Tier 4 | Nothing — but the record often lives only in the algo governance file, not the AI inventory |
| Collections, arrears and forbearance prioritization | D1 3–4 · D2 2 · D3 3 · D4 3 · D5 3 | Tier 3 | Vulnerability is treated as a data attribute rather than a consequence multiplier on D1 |
| Retail customer assistant answering product questions | D1 3 · D2 2–3 · D3 3 · D4 4 · D5 3 | Tier 3 | D3 scored 1 because the chat log can be deleted; the customer has already read it |
| Reconciliation or journal-posting agent | D1 3 · D2 4 · D3 3 · D4 3 · D5 2 | Tier 4 | See the worked dispute in Chapter 7, Example D |
| KYC document extraction feeding a human reviewer | D1 2–3 · D2 2 · D3 2 · D4 3 · D5 4 | Tier 2–3 | D5 under-read: identity documents are among the most sensitive inputs the firm holds |
| Analyst research summarization on internal documents | D1 2 · D2 1 · D3 2 · D4 2 · D5 3–4 | Tier 2 | Material non-public information in the grounding corpus, scored as ordinary confidentiality |
| Drafting of model documentation or validation evidence | D1 3 · D2 1 · D3 2 · D4 2 · D5 3 | Tier 2, with a control addition | Treated as a productivity tool although its output is the firm's assurance record |
2. The regulatory interface
Regulatory note. The regimes below are named so each IRGF record can be pointed at the obligation it evidences, not to restate them. Applicability, thresholds and commencement dates differ by jurisdiction and several have moved during implementation. Nothing here is legal advice: confirm the current position with your own counsel, and record the answer in the Regulatory Overlay Reference so it is checkable later.
IRGF does not restate any of these obligations. It gives each one a record that carries the evidence, an owner, and a trigger that reopens it when the obligation or the system changes.
| Regime or standard | What it obliges in practice | IRGF record that carries the evidence |
|---|---|---|
| Model risk management supervision (for example SR 11-7 / OCC 2011-12 in the US, PRA SS1/23 in the UK) | A model inventory, tiering, independent validation proportionate to tier, ongoing monitoring and a documented owner. Most AI systems in scope of IRGF are also models under these definitions. | Risk Classification Record reconciled with the model inventory; validation output referenced from the AI Assurance Summary rather than repeated |
| Automated decision-making and profiling rules (for example GDPR Article 22 and equivalents) | A meaningful human oversight point, an explanation of the logic involved, and a route to contest a decision. | Human oversight point confirmed at G3 in the Deployment Authorization Record; D2 anchor evidence from override-rate instrumentation |
| Adverse action and fair lending obligations (for example ECOA/Regulation B, FCRA, EBA loan origination guidelines) | Specific, accurate reasons for a decline; no reliance on a reason the firm cannot substantiate; monitoring of outcomes. | Reason-code generation recorded as a control in the Control Matrix, mapped to the D5 uncertainty score |
| EU AI Act, where it applies | Creditworthiness assessment of natural persons sits in the Annex III high-risk list, with obligations on risk management, data governance, logging, human oversight, accuracy and post-market monitoring. Life and health insurance pricing is treated similarly. | Regulatory Overlay Reference mapping each affected system; logging obligations satisfied from the audit link of the agent chain |
| Digital operational resilience rules (for example DORA in the EU) | ICT third-party register, concentration risk assessment, exit strategies, incident classification and reporting, resilience testing. | Vendor fields of the AI-extended ADR; Drift/Alert Record routing into the existing incident taxonomy |
| Algorithmic trading requirements (for example MiFID II RTS 6) | Pre-deployment testing in a segregated environment, kill functionality, annual self-assessment, documented change control. | Kill-switch conditions pre-approved at G3; annual self-assessment aligned to the G4 re-authorization cycle rather than run as a separate exercise |
| Risk data aggregation principles (BCBS 239) | Accuracy, completeness and traceability of the data feeding risk reporting, including lineage. | Data Lineage and Sensitivity Record, which is the same artifact the grounding sources need anyway |
| Conduct and fair-treatment regimes (for example FCA Consumer Duty) | Evidence that communications support understanding, that outcomes are monitored, and that foreseeable harm is avoided — including for customers in vulnerable circumstances. | Outcome monitoring registered as a runtime signal; vulnerability handling recorded as a D1 justification rather than a data field |
3. Calibrating the five dimensions
The dimensions do not change. What changes is what a 3 and a 4 look like when the subject matter is this sector, and which reading an assessor under delivery pressure reaches for first.
| Dimension | How to read it here | The mis-score to watch for |
|---|---|---|
| D1 Decision Consequence | Any output that determines access to credit, to an account, or to a customer's own funds starts at 3. Denial, withdrawal or closure, and anything that produces an adverse record a third party will rely on, is 4. | Scoring on the firm's loss rather than the customer's. A decision that costs the bank nothing and the customer their tenancy application is a 4. |
| D2 Autonomy | Operational review under volume targets is the sector's characteristic autonomy trap. A credit officer processing 300 proposals a day is not exercising judgment on each. | Recording 2 for “human approves each” without measuring the override rate. Instrument it before the score is countersigned, not after. |
| D3 Reversibility Deficit | Money that has moved, a report that was not filed in its window, a bureau record, and a statement made to a customer are all effectively irreversible regardless of what the ledger can undo. | Mechanism reversibility. A reversal entry restores the balance; it does not un-file the regulatory return that used the wrong number. |
| D4 Exposure and Scale | Count the reach of the records the system writes into, not the number of people who log in. A system feeding the general ledger or a regulatory return has enterprise reach with four users. | Scoring 2 for “internal, finance team only” when the output lands in the statutory accounts. |
| D5 Sensitivity and Uncertainty | Financial transaction data, identity documents and material non-public information each sit at 3 or above. Take the higher of the sensitivity and uncertainty sub-scores; do not average them. | Treating MNPI as ordinary confidential data because the classification scheme has no category for it. |
4. One system, end to end
The overlay above is a map. This is one route across it: a single system, followed from the moment data arrives to the moment the consequence leaves the firm, with the record that attaches at each step.
5. What each gate adds
Additions only. Everything in the base gate definitions still applies; see the gate checklists for the common set.
| Gate | Sector addition | Why it is here |
|---|---|---|
| G1 | State whether the system is a model under the firm's existing model risk policy, and record the answer. State whether the grounding corpus can contain material non-public information. | Both answers determine which existing process owns the work. Deciding late means either duplicate review or an uncovered system. |
| G2 | Lineage to BCBS 239 standard for anything feeding risk or regulatory reporting. Vendor concentration assessment where the model, the hosting and the retrieval layer come from one supplier. | Concentration is invisible at system level and only appears when the portfolio is looked at once. |
| G3 | Independent validation sign-off referenced, not repeated. Reason-code coverage evidenced for any declining decision. Kill-switch conditions for anything acting on payments or orders. | Validation answers “is the model sound”. G3 answers “is this safe to run here, now, at this tier”. Collapsing them loses the second question. |
| G4 | Any vendor model version change at Tier 3–4 is Material by default. Challenger promotion is a Major change, not a tuning exercise. | The most common uncontrolled change in banking AI is a supplier upgrading a model under a managed service. |
| G5 | Retention set by the longest applicable record-keeping obligation, and evidence retrievable in a form a regulator can read after the platform is gone. | Retirement usually removes the ability to answer questions about decisions the system made while it ran. |
6. Controls and evidence worth adding
| Control | Where it attaches | Evidence it produces |
|---|---|---|
| Single AI and model inventory, reconciled monthly | Risk Classification Record, model inventory | A reconciliation report showing zero systems in one inventory and not the other |
| Override-rate instrumentation on every D2 = 1 or 2 system with D1 ≥ 3 | Runtime telemetry, reviewed at the first assurance cycle | A measured override rate against a threshold agreed in advance |
| Reason-code generation and substantiation check | Control Matrix, mapped to the adverse action obligation | Sampled decisions with the reason given and the factor that actually drove it |
| Value, counterparty and rate limits for any acting agent | Agent Card, enforced in the policy engine | Machine-readable boundary, comparable against the IAM grant |
| Segregation of material non-public information from retrieval corpora | Data Lineage and Sensitivity Record; retrieval index build | Source-level classification and an index build log that excludes restricted classes |
| Dual control on any change to an agent's authority boundary | Change record, policy repository | Two named approvers and a diff of the boundary |
| Alert suppression rate monitoring for transaction monitoring | Runtime signal into the Control Plane | Suppression rate by typology, with a change threshold that triggers re-score |
7. Runtime signals to wire first
Control Plane onboarding order matters more than coverage in the first year (Chapter 18). These are the signals that earn their place earliest in this sector.
| Signal | Drift category | Suggested response |
|---|---|---|
| Override rate on a reviewed decision system falls below the agreed floor | Behavioral, with a classification consequence | Re-score D2 rather than note it; the recorded tier is wrong, not the behaviour |
| Supplier changes the underlying model version of a managed service | AI-model behavioral | Tier 3–4: treat as Material change and route to G4. Tier 1–2: log and batch-review |
| Retrieval corpus gains a source above its approved sensitivity | Data lineage and governance | Block the index build if the pipeline can enforce it; otherwise same-day owner alert and re-score of D5 |
| Agent posting outside its approved account range or above its value limit | Agent boundary divergence | Incident at any tier: suspend, revoke, preserve state, notify the Agent Owner and AI Governance Lead |
| Alert suppression rate moves more than the agreed band, in either direction | Behavioral | Route to the financial crime team with the model owner, not to the architect |
| Consumer or conduct rules change in a mapped jurisdiction | Policy and regulatory | Re-open compliance assurance for every system in the overlay mapping, without waiting for a release |
8. Failure modes this sector produces
Two inventories, two review queues, one system
The tell. Model risk asks for a validation pack and AI governance asks for an assurance summary, and the delivery team writes the same content twice with different words. Within two quarters the two records disagree and nobody knows which is authoritative.
The response. Decide the interface once: the model inventory is the register, the Risk Classification Record is the tiering, and the assurance summary references the validation report rather than restating it. Reconcile monthly and publish the exceptions.
The assistant that quietly became a decision system
The tell. A generative assistant deployed to “help advisers draft responses” is, six months later, producing the response that goes out with a one-click send. Nothing was reclassified because nothing was rebuilt.
The response. Autonomy change is a change type in its own right (Chapter 17). Detect it from the interaction telemetry — the ratio of edits to sends — not from the change log.
Validation evidence written by the system under validation
The tell. An LLM drafts the model documentation, the validation summary and the challenge log. The reviewer reads fluent text that agrees with itself and signs.
The response. Permit drafting; prohibit the drafting system from being the source of the evidence. Require the challenge log to record at least one substantive disagreement and its resolution, authored by a named person.
Tier laundering through the human-in-the-loop claim
The tell. Every high-consequence system arrives at G2 with D2 = 2 and a screenshot of an approval button.
The response. Countersignature is the control (Chapter 7, §7.6), and the override-rate threshold must be agreed at G2, before anyone knows the answer.
Concentration discovered during an outage
The tell. Model, embeddings, vector store and the assistant framework all come from one supplier, and each system cleared G2 on its own merits.
The response. Concentration is a portfolio property. Add a standing quarterly question to the Architecture Review Board: which single supplier failure would take out the most Tier 3–4 systems, and what is the exit?
9. A ninety-day start
If the sector is yours and the framework is new, this is the order that produces something defensible fastest. It assumes one part-time architect and one risk lead, not a programme.
- List every AI system touching credit, payments, financial crime and customer communications. Include the ones procured as features of existing platforms — that is where the omissions are.
- Reconcile that list against the model inventory. Publish the two exception lists: models that are not in the AI inventory, and AI systems nobody has tiered.
- Agree the interface with model risk in one page: who tiers, who validates, which record is authoritative, and what G3 adds that validation does not.
- Score the five highest-consequence systems with an independent countersigner present, using the classification checklist.
- Instrument override rate on the two reviewed decision systems with the highest volume. Agree the floor before you see the number.
- Write authority boundaries for every agent that touches a ledger, a payment rail or an order, using the Agent Card. Boundaries first, kill-switches second.
- Wire two drift signals only: grounding-source change on the highest-tier retrieval system, and supplier model version change across the estate.
- Run one G3 dry run on a system already in production. The gap list it produces is your real backlog.
The order matters. Firms that begin with tooling spend the first year integrating a repository and the second discovering that the tiers in it were never defensible.