AI agents in banking · The build map · Brief 06

Model risk management: a design brief for banks.

Last updated Sep 17, 2026 · standard answers for this use case · adjust them in the map

Model risk teams are a natural early adopter because their output is documents: development records, validation reports, monitoring summaries and the inventory itself. A chained workflow that drafts each artefact from the underlying evidence and checks it against the standard the team already applies removes weeks of writing. The line to hold is that the language model drafts and checks; the validator's judgment and signature remain the validator's.

WorkflowPrompt chainingEvaluator and optimizerTier 3: act within an envelopeautonomy level 3

Model risk management: The model may act within a defined envelope; people sample, monitor and can stop it. Several model calls orchestrated by your code, in a sequence or a graph you designed. Low stakes, reversible, internal. The model may act within a defined envelope, with sampling, monitoring and a kill switch instead of per-action review.

PatternWorkflow. Your code lays out the steps, the cost of error is real, and you need to observe every step.Prompt chainingEach step's output is the next step's input, with a programmatic check between them.Evaluator and optimizerOne model produces, another checks against explicit criteria, and the loop repeats until the check passes or a person is called.KnowledgeRetrieval over the governed document set, with a citation on every answer and a refusal when nothing relevant is found.Delivery routeProvider API: one third party to diligence; residency and retention terms are yours to negotiate.

Which steps belong to a person, the model and a system?

A PERSONTHE MODELA SYSTEM1
Assemble the evidence
2
Draft the validation report sections
3
Check against the validation standard
4
Validate and sign
5
Update the inventory
#StepOwnerNote
1Assemble the evidenceA systemModel code, data lineage, test results from the model repository.
2Draft the validation report sectionsThe modelConceptual soundness, monitoring, outcomes analysis, each cited.
3Check against the validation standardThe modelAn evaluator pass against the bank's own checklist.
4Validate and signA personEffective challenge is a person's job.
5Update the inventoryA systemStructured fields written by code from the signed report.

Which control layers carry the weight?

Governance and accountabilityCOREIdentity and entitlementsLIGHTAction gatewayLIGHTData and knowledgeCOREModels and vendorsLIGHTRuntime and orchestrationSTANDARDObservability, evaluation and auditSTANDARDHuman oversight and escalationSTANDARD

Each layer is described, with its controls and documents, on the control plane page.

Which rules and guidance does this design answer to?

DocumentAuthorityWhy it applies hereStatus
SR 26-2Federal ReserveThe 2026 US framework: narrower model definition, materiality, generative and agentic AI carved out.In force
OCC Bulletin 2026-13OCCSame text for national banks; the promised interagency RFI on AI.In force
PRA SS1/23UK (BoE / PRA / FCA)UK: AI and machine learning stay inside model risk management.In force
ECB Guide to internal models (July 2025, ML section)ECBEU capital models: explainability and justified complexity.In force
SR 23-4Federal ReserveUS third-party risk management, including the model provider.In force
SB 26-189Colorado AI ActColorado: notice, explanation and human review for consequential automated decisions from Jan 1, 2027.Final
FSB AI monitoring report (Oct 2025)FSBConcentration on a small number of model suppliers.Final
NIST AI RMF 1.0NISTThe voluntary Govern, Map, Measure, Manage frame for everything model-risk guidance leaves out.In force

How will you know it works, before and after launch?

CODE CHECKS30%JUDGE MODEL50%HUMAN REVIEW20%
  • A golden dataset of at least 50 real cases with expected outputs, including adversarial inputs: wrong documents, unusual formats, prompts that try to change the task.
  • A judge model scoring against a written rubric (accuracy, completeness, tone, citation present), calibrated against a human-scored sample every month. Threshold set from the human sample, not guessed.
  • Citation checks: every factual claim resolves to a passage in the governed set; unsupported claims below 2% of answers.
  • Trace evals per step, not only end to end: which step fails, how often, at what cost, so a prompt or model change can be judged step by step.
  • The same suite reruns on every prompt change, model version change and retrieval change; a regression blocks the release. That is what ongoing monitoring and outcomes analysis mean in model-risk terms.

Where must a person be in the loop?

  • An action envelope: which tools, which systems, what value, what volume, what the model may never do.
  • Sampled human review of outputs and a weekly look at the exception log.
  • Turn and cost budgets, logged, with an automatic stop when exceeded.
  • A kill switch that any owner can pull.
  • Drafts that read as validation without the tests having run.
  • Using the tool on generative systems and calling the result model validation; the 2026 guidance says it is not.
  • No version control on the prompts that produce examinable documents.

What will a validator or an examiner ask?

  1. Where is this system in your inventory, what tier did you assign, and who signed it off?
  2. What counts as a model here, and what does your validation cover for the parts that are not?
  3. Show me the data lineage behind the retrieval set and the training or tuning data.
  4. What can the system do without a person, and where is that written down?
  5. How do you know it is still working: which evals run, how often, and what happened the last time one failed?
  6. What did you do about the vendor: due diligence, contract, exit plan, concentration?
  7. Walk me through one wrong output from production and what the customer, if any, saw.
  8. Who can switch it off, and has that been tested?

What it does: documents, validates and monitors models, and drafts the artefacts examiners read; the model may act within a defined envelope; people sample, monitor and can stop it.

Pattern: workflow (prompt chaining, evaluator and optimizer); tier 3: act within an envelope.

Rules it answers to: 4 documents across US, each linked in the brief.

How we know it works: a golden dataset, automated checks on every release, and human review at the level the tier demands.

What could go wrong and who answers: the accountable owner, the kill switch, the escalation route.

Which of the 100 largest US banks have put this use case on the record?

Every rule this brief cites, the morning it changes.

agent deployments, regulator positions and the day's six stories · in your inbox by 7 am ET · free

plus every tracker, bank and agent page update, the morning after · leave any morning