AI agents in banking · The build map · Brief 07

Generative & agentic AI: a design brief for banks.

Last updated Sep 17, 2026 · standard answers for this use case · adjust them in the map

An agent is the right pattern only when the path cannot be fixed in advance and the stakes allow the model to choose it. Inside a bank that means internal, reversible work: research across documents and systems, drafting with tool access, working an exception queue. The design that supervisors have described approvingly is a bounded one: a small, named tool set, a turn and cost budget, an action envelope, full traces, and a person sampling the output. The control plane and lifecycle on this site were written for exactly this case.

Bounded agentTier 3: act within an envelopeautonomy level 3

Generative & agentic AI: The model may act within a defined envelope; people sample, monitor and can stop it. The model plans its own path through a set of tools toward a goal, inside a turn budget and an action envelope. Low stakes, reversible, internal. The model may act within a defined envelope, with sampling, monitoring and a kill switch instead of per-action review.

PatternBounded agent. The path cannot be fixed in advance and the stakes allow it. Autonomy is kept in check by limiting tools, turns and what an action may touch.KnowledgeRetrieval over the governed document set, with a citation on every answer and a refusal when nothing relevant is found. Tool calls to the system of record for anything live. Never retrieval for a balance, a case status or a limit. Documents to reason with, systems to check against: the workflow decides which question goes where before the model answers.Delivery routeThrough a cloud platform: two third parties in the chain and a concentration question; often the faster route through security review.

Which steps belong to a person, the model and a system?

A PERSONTHE MODELA SYSTEM1
Set the goal and the envelope
2
Plan and act
3
Enforce the envelope
4
Evaluate the result
5
Sample and approve
#StepOwnerNote
1Set the goal and the envelopeA personWhat done looks like, which tools, what the agent may never do.
2Plan and actThe modelTool calls within the envelope; every call traced.
3Enforce the envelopeA systemThe action gateway checks identity, entitlements, value and volume before any call executes.
4Evaluate the resultThe modelAn evaluator pass against the goal; failures stop or escalate.
5Sample and approveA personSampled review; per-action approval for anything that leaves the bank.

Which control layers carry the weight?

Governance and accountabilityCOREIdentity and entitlementsCOREAction gatewayCOREData and knowledgeCOREModels and vendorsCORERuntime and orchestrationCOREObservability, evaluation and auditCOREHuman oversight and escalationSTANDARD

Each layer is described, with its controls and documents, on the control plane page.

Which rules and guidance does this design answer to?

DocumentAuthorityWhy it applies hereStatus
SR 26-2Federal ReserveGenerative and agentic AI are outside model-risk scope; broader governance applies.In force
OCC Semiannual Risk Perspective, Spring 2026OCCSupervisors observed 'measured' agentic adoption: guardrails and human-in-the-loop accountability.Final
FSB AI sound practices consultation (June 2026)FSBTwelve sound practices, with specific attention to generative and agentic AI.Proposed
CAISI RFI on AI agent security (2026)NISTThe security questions being asked of AI agents.Proposed
NIST AI 600-1 (Generative AI Profile)NISTTwelve generative-AI risks and 200-plus actions.In force
SR 23-4Federal ReserveUS third-party risk management, including the model provider.In force
SB 26-189Colorado AI ActColorado: notice, explanation and human review for consequential automated decisions from Jan 1, 2027.Final
BCBS Third-Party Risk Principles (Dec 2025)Basel CommitteeNth-party supply chains and concentration on cloud providers.In force
BCBS ICT Risk Management Report (June 2026)Basel CommitteeHow supervisors look at ICT and cloud dependencies.Final
NIST AI RMF 1.0NISTThe voluntary Govern, Map, Measure, Manage frame for everything model-risk guidance leaves out.In force

How will you know it works, before and after launch?

CODE CHECKS30%JUDGE MODEL50%HUMAN REVIEW20%
  • A golden dataset of at least 50 real cases with expected outputs, including adversarial inputs: wrong documents, unusual formats, prompts that try to change the task.
  • A judge model scoring against a written rubric (accuracy, completeness, tone, citation present), calibrated against a human-scored sample every month. Threshold set from the human sample, not guessed.
  • Citation checks: every factual claim resolves to a passage in the governed set; unsupported claims below 2% of answers.
  • Trace evals per step, not only end to end: which step fails, how often, at what cost, so a prompt or model change can be judged step by step.
  • The same suite reruns on every prompt change, model version change and retrieval change; a regression blocks the release. That is what ongoing monitoring and outcomes analysis mean in model-risk terms.

Where must a person be in the loop?

  • An action envelope: which tools, which systems, what value, what volume, what the model may never do.
  • Sampled human review of outputs and a weekly look at the exception log.
  • Turn and cost budgets, logged, with an automatic stop when exceeded.
  • A kill switch that any owner can pull.
  • Unbounded autonomy: no turn budget, no envelope, no kill switch.
  • Granting the agent the operator's own identity instead of its own.
  • Assuming model-risk validation covers it; in the US it does not, so governance must.

What will a validator or an examiner ask?

  1. Where is this system in your inventory, what tier did you assign, and who signed it off?
  2. What counts as a model here, and what does your validation cover for the parts that are not?
  3. Show me the data lineage behind the retrieval set and the training or tuning data.
  4. What can the system do without a person, and where is that written down?
  5. How do you know it is still working: which evals run, how often, and what happened the last time one failed?
  6. What did you do about the vendor: due diligence, contract, exit plan, concentration?
  7. Walk me through one wrong output from production and what the customer, if any, saw.
  8. Who can switch it off, and has that been tested?

What it does: runs a bounded agent on internal tasks: research, drafting, operations exceptions; the model may act within a defined envelope; people sample, monitor and can stop it.

Pattern: bounded agent; tier 3: act within an envelope.

Rules it answers to: 4 documents across US, each linked in the brief.

How we know it works: a golden dataset, automated checks on every release, and human review at the level the tier demands.

What could go wrong and who answers: the accountable owner, the kill switch, the escalation route.

Which of the 100 largest US banks have put this use case on the record?

Every rule this brief cites, the morning it changes.

agent deployments, regulator positions and the day's six stories · in your inbox by 7 am ET · free

plus every tracker, bank and agent page update, the morning after · leave any morning