Model governance and AI-assisted risk work
Keep model use bounded, reviewable, and accountable.
A model can be wrong in a sophisticated way. An AI assistant can write a fluent case note with an invented fact. Governance makes those failures visible before the output becomes an unchecked decision.
Use the current model-risk framework
The federal banking agencies issued revised model-risk guidance in April 2026 through SR 26-2, replacing SR 11-7 and SR 21-8. The letter states its expected relevance to Federal Reserve-regulated banking organizations over $30 billion in assets and emphasizes a risk-based approach. Applicability must be assessed for the actual institution.
For this textbook’s engineering design, maintain an inventory of models and their uses, owners, limitations, dependencies, and review evidence. A small rule-based model can still be important if it controls a large exposure. Governance effort should reflect the consequence and complexity of use rather than the prestige of the algorithm.
Inside the mechanism. Governance should reflect the actual model use, institution, materiality, and applicable framework. Inventory models and material decision components by function rather than relying on a team’s preferred label. Record owner, purpose, inputs, dependencies, limitations, validation status, and approved uses. Current guidance must be checked for scope and effective changes before treating an older supervisory document as the governing reference.
A concrete example. Governance begins with the actual model use, materiality, institution, and applicable supervisory framework. An inventory entry needs a decision owner and current evidence. The case identifies 1,428 eligible records from a source population of 1,700. The required workflow completes for 1,385, but 21 completed records miss the illustrative internal target. Another 43 remain incomplete. Communication evidence covers 1,371 generated notices. Scope, completion, timeliness, and delivery are four separate properties of the customer outcome.
When the assumption fails. A material decision model is omitted because the team calls it a rule or vendor score. Inventory decision components by function and apply the relevant review and change process. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.
Governance begins with the actual model use, materiality, institution, and applicable supervisory framework. An inventory entry needs a decision owner and current evidence.
- InventoryIdentify models and actual uses
- AssessEvaluate consequence complexity and limitations
- GovernAssign proportionate review and ownership
- Algorithm complexity
- Technical sophistication of the method
- Use risk
- Consequence of relying on its output
Model inventory
Illustrative data; not a real customer record or a prescribed policy.
- Modelsimple reserve estimator
Low algorithm complexity
- Exposurelarge merchant portfolio
High consequence
- Reviewproportionate to use
Complexity alone is insufficient
A simple model can support a consequential decision
Assess model risk in its actual use. A simple model can support a consequential decision.
- Failure mode 1avoid
- Treat SR 11-7 as the current sole guidance. The 2026 letter replaced it.
- Failure mode 2avoid
- Apply the banking letter identically to every startup. Institutional scope matters.
- Failure mode 3avoid
- Govern only machine-learning models. Other quantitative tools can be important.
Separate development from effective challenge
Developers understand the model deeply, but they also know the assumptions they intended. Independent challenge tests whether those assumptions hold and whether the use is appropriate. The structure should fit the organization and applicable expectations.
Review conceptual soundness, data, implementation, outcomes, limitations, and controls. Track findings through remediation and retest. A reviewer’s signature is not evidence that every concern was resolved. The model owner should know which limitations remain and what use is permitted while they remain.
Effective challenge examines whether the model is suitable for its stated use, including limitations that the development team may not have emphasized. It can inspect the target, data, assumptions, validation evidence, implementation, and downstream policy. Independence is useful because the people responsible for delivery may face pressure to interpret ambiguous results favorably. The form and depth of review should fit the model’s materiality and the applicable framework.
A limitation becomes operational when it changes the permitted use. If a model has weak evidence for a new product segment, the response may be restricted deployment, more review, additional validation, or another approved control. Recording the limitation in a document while using the model without restriction does not manage it.
Inside the mechanism. Effective challenge needs access to the relevant data, methods, assumptions, and outcomes, plus authority to require a response. The reviewer should test what could invalidate the development conclusion. Keep findings, management responses, unresolved limitations, and release conditions explicit. Organizational separation alone is not evidence of a substantive review, and a repeated developer summary is not independent validation.
A concrete example. A reviewer needs sufficient evidence, expertise, and authority to question the development conclusion. Repeating the developer summary does not test its assumptions. The case identifies 673 eligible records from a source population of 740. The required workflow completes for 653, but 10 completed records miss the illustrative internal target. Another 20 remain incomplete. Communication evidence covers 646 generated notices. Scope, completion, timeliness, and delivery are four separate properties of the customer outcome.
When the assumption fails. Validation signs off without access to the evaluation population or known limitations. Document challenges, test evidence, unresolved limitations, and release conditions. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.
A reviewer needs sufficient evidence, expertise, and authority to question the development conclusion. Repeating the developer summary does not test its assumptions.
- DevelopDocument assumptions and evidence
- ChallengeTest the design and actual use
- ResolveTrack findings and permitted limitations
- Developer explanation
- Why the model was built this way
- Independent challenge
- Evidence that tests the explanation
Review finding
Illustrative data; not a real customer record or a prescribed policy.
- Issueweak new-merchant performance
Identified limitation
- Use restrictionknown merchants only
Bounded permitted use
- Retestscheduled with new evidence
Finding remains tracked
Approval alone does not remove limitations
Connect review findings to use restrictions and retesting. Approval alone does not remove limitations.
- Failure mode 1avoid
- Treat a signature as full proof. The underlying findings matter.
- Failure mode 2avoid
- Let unresolved issues disappear at launch. They remain part of the risk.
- Failure mode 3avoid
- Review only code style. Conceptual and data weaknesses can dominate.
Bound generative AI to evidence-supported tasks
Generative AI can help summarize records, draft narratives, or retrieve relevant policy passages. It can also invent facts, omit context, or follow instructions embedded in untrusted documents. Treat retrieved customer material as evidence to analyze, not instructions to the system.
Use constrained inputs, source references, and a reviewable output. Require factual claims to point to supporting records. Keep sensitive reporting information within approved access boundaries. A model-generated narrative should not automatically file a report, release funds, or change a customer restriction without the authorized decision process.
A generative system can help summarize retained evidence, but fluent wording is not evidence. Each material factual statement should be traceable to the permitted source material, and the workflow needs a response when support is absent or contradictory. Treat retrieved documents and customer submissions as data rather than trusted instructions. Separate the assistant’s proposal from the authorized decision, especially when the action changes funds, account access, or a regulated process. Evaluation should include fabricated facts, missing evidence, private-data exposure, and misleading instructions embedded in source material.
Inside the mechanism. Bound an AI assistant by task and authority. A case-summary tool should distinguish quoted source facts, supported synthesis, and unresolved questions, and link material claims to evidence. It should not acquire the authority to move money or make a legal determination merely because it can produce fluent text. Treat retrieved documents and customer messages as data that can contain misleading instructions. Preserve the original sources for review.
A concrete example. An assistant can summarize a case while still introducing unsupported facts or omitting decisive evidence. Its output needs links to the source record and a bounded role. The rule flags 409 of 6,200 reviewed summaries. Of those flags, 322 meet the synthetic target, giving 78.73% precision. It misses 81 target events. Under the stated cost assumptions, residual loss and operating friction total $9,361. The important result is the connection between the population, action, capacity, and outcome—not one isolated score.
When the assumption fails. Generated text presents an unverified suspicion as an established fact. Require evidence references, explicit uncertainty, and authorized human decisions for the defined task. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.
An assistant can summarize a case while still introducing unsupported facts or omitting decisive evidence. Its output needs links to the source record and a bounded role.
- RetrieveProvide approved relevant evidence
- DraftGenerate a bounded supported output
- ReviewVerify claims before consequential action
- Fluent narrative
- Text reads plausibly
- Grounded narrative
- Claims are supported by the supplied records
AI drafting record
Illustrative data; not a real customer record or a prescribed policy.
- Claimcustomer admitted intent
Generated statement
- Sourcenone
Unsupported claim
- Actionremove and investigate
Do not accept fluent invention
Fluency does not establish truth
Require claim-level evidence and authorized review. Fluency does not establish truth.
- Failure mode 1avoid
- Let customer documents instruct the assistant. They are untrusted task data.
- Failure mode 2avoid
- Auto-release funds from a generated summary. Consequential authority needs controls.
- Failure mode 3avoid
- Copy restricted reports into broad AI tools. Access and data rules still apply.
Evaluate AI failure modes before use
An AI evaluation should reflect the actual task: factual accuracy, omissions, unsupported claims, confidentiality, instruction handling, and consistency. Include difficult cases and deliberate misleading content in controlled fixtures.
Measure the harm of errors, not only an average quality score. A rare invented admission can be more consequential than several awkward sentences. Compare against a useful baseline and define when a human must intervene. Version prompts, retrieval rules, models, and evaluation sets so a vendor update does not silently invalidate the evidence.
Inside the mechanism. Evaluate failures that matter to the task: unsupported claims, omitted decisive facts, wrong entities, incorrect amounts, confidentiality errors, and susceptibility to instructions embedded in evidence. Use representative records and targeted difficult cases with a stated scoring rubric. Measure critical errors separately from style. A high average score can hide an unacceptable failure in a small but consequential class of cases.
A concrete example. Average fluency does not measure unsupported claims, instruction misuse, missing evidence, or sensitive-data exposure. Evaluation needs a labeled failure taxonomy tied to the use case. The rule flags 339 of 3,100 evaluation records. Of those flags, 298 meet the synthetic target, giving 87.91% precision. It misses 74 target events. Under the stated cost assumptions, residual loss and operating friction total $10,811. The important result is the connection between the population, action, capacity, and outcome—not one isolated score.
When the assumption fails. A polished summary passes review while missing the fact that changes the case disposition. Test critical omissions and unsupported claims across representative and adversarial records. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.
Average fluency does not measure unsupported claims, instruction misuse, missing evidence, or sensitive-data exposure. Evaluation needs a labeled failure taxonomy tied to the use case.
- TaskDefine the permitted output and harm model
- EvaluateTest representative and adversarial fixtures
- GateUse thresholds and human review appropriate to consequence
- Average fluency
- General writing quality
- Critical error rate
- Frequency of consequential unsupported or unsafe output
AI evaluation
Illustrative data; not a real customer record or a prescribed policy.
- Cases200 synthetic records
Controlled test set
- Invented facts3
Critical errors
- Decisionnot ready for autonomous use
Average style score is insufficient
A good average can hide unacceptable rare failures
Evaluate consequential errors separately. A good average can hide unacceptable rare failures.
- Failure mode 1avoid
- Score only writing style. Truth and confidentiality remain untested.
- Failure mode 2avoid
- Reuse old evaluations after major changes. The system behavior may differ.
- Failure mode 3avoid
- Use real confidential cases without approval. Test data also has access constraints.
Maintain a complete change record
Model behavior can change through data, features, code, thresholds, prompts, retrieval sources, or vendor versions. Keep a release record that connects the change to evaluation, approval, rollout, monitoring, and rollback.
Define what counts as a material change and who decides. A prompt edit that adds a new tool can be more consequential than a model patch. Preserve the affected population and outputs for review under the retention policy. Governance is complete when the organization can explain what changed, why it was allowed, and how the result was checked.
Inside the mechanism. A complete change record includes data, transformations, model or prompt version, tools, retrieval sources, thresholds, and downstream policy. Hosted vendor changes can alter behavior without a local code change. Record available version controls and evaluate material changes within the approved process. Rollback needs a known configuration and a way to identify decisions made during the affected interval.
A concrete example. A model change can include data, features, prompts, vendor versions, thresholds, and downstream policy. The release record should identify the full decision configuration. The case identifies 874 eligible records from a source population of 930. The required workflow completes for 848, but 13 completed records miss the illustrative internal target. Another 26 remain incomplete. Communication evidence covers 840 generated notices. Scope, completion, timeliness, and delivery are four separate properties of the customer outcome.
When the assumption fails. A vendor changes a hosted model without a corresponding internal review record. Pin available versions, assess material changes, and preserve approval and rollback evidence. The following worked sequence shows the reference condition, a stress condition, and a response condition with explicit synthetic data. These are comparative assumptions, not measured causal effects.
A model change can include data, features, prompts, vendor versions, thresholds, and downstream policy. The release record should identify the full decision configuration.
- ChangeIdentify all behavior-affecting components
- EvidenceAttach evaluation and approval
- OperationMonitor rollout and preserve rollback
- Code version
- One implementation component
- System version
- Model data policy prompts and dependencies together
AI release record
Illustrative data; not a real customer record or a prescribed policy.
- Promptv5
Changed instructions
- Modelprovider version recorded
Dependency identity
- Toolsread-only evidence search
Bounded capability
Behavior can change outside application code
Version the whole decision system. Behavior can change outside application code.
- Failure mode 1avoid
- Review only model weights. Prompts and data can alter outcomes.
- Failure mode 2avoid
- Treat tool access as a minor wording edit. Capabilities change the risk.
- Failure mode 3avoid
- Omit rollback and monitoring. The release lacks an operational response.
Chapter connections
This chapter builds on Experiments, causal effects, and risk tradeoffs. Use the glossary for terminology and risk mathematics for formulas and worked calculations.