AI agents in banking · The build map · Brief 10

Data & privacy: a design brief for banks.

Last updated Sep 17, 2026 · standard answers for this use case · adjust them in the map

Data work is where the older standards bite hardest on the newest systems: a training set, a feature store and a retrieval corpus are risk data under BCBS 239, and personal data under privacy law. The useful patterns are chained and code-checked: classify documents and fields against a schema, propose lineage and ownership records, draft data-protection assessments and handle access requests, with every output validated by rules and a person accountable for the register.

Augmented callTier 2: act with approvalautonomy level 1

Data & privacy: The model prepares and may act within limits; a person approves anything that reaches a customer or cannot be undone. One model call with retrieval, tools and a structured output, inside code you control. Material but recoverable. The model may prepare and, within limits, act, with a person approving anything that leaves the bank or touches a customer.

PatternAugmented call. The task is well defined and the output can be checked. Most bank use cases should start here.KnowledgeRetrieval over the governed document set, with a citation on every answer and a refusal when nothing relevant is found. Tool calls to the system of record for anything live. Never retrieval for a balance, a case status or a limit. Documents to reason with, systems to check against: the workflow decides which question goes where before the model answers.Delivery routeProvider API: one third party to diligence; residency and retention terms are yours to negotiate.

Which steps belong to a person, the model and a system?

A PERSONTHE MODELA SYSTEM1
Discover and classify
2
Validate classifications
3
Record lineage and ownership
4
Draft assessments and responses
5
Own the register
#StepOwnerNote
1Discover and classifyThe modelData types and sensitivity against the bank's schema, with a confidence.
2Validate classificationsA systemRules and sampling; low-confidence items to a person.
3Record lineage and ownershipA systemStructured writes to the catalogue from validated results.
4Draft assessments and responsesThe modelImpact assessments, access-request replies, from the register.
5Own the registerA personThe data owner signs; the privacy officer reviews assessments.

Which control layers carry the weight?

Governance and accountabilityCOREIdentity and entitlementsCOREAction gatewayCOREData and knowledgeCOREModels and vendorsLIGHTRuntime and orchestrationLIGHTObservability, evaluation and auditSTANDARDHuman oversight and escalationSTANDARD

Each layer is described, with its controls and documents, on the control plane page.

Which rules and guidance does this design answer to?

DocumentAuthorityWhy it applies hereStatus
BCBS 239Basel CommitteeRisk data must be owned, traceable, complete and current; training data is risk data.In force
Regulation (EU) 2024/1689EU AI ActArticle 10: documented data governance for high-risk systems.In force
GDPR Article 22EU AI ActRights around solely automated decisions with legal or similar effect.In force
CPPA ADMT, risk-assessment and cybersecurity-audit regulationsCalifornia CPPACalifornia: notice, opt-out and risk assessments for automated decision-making technology.In force
SR 26-2Federal ReserveUS model risk management.In force
PRA SS1/23UK (BoE / PRA / FCA)UK: the widest model perimeter.In force
FSB AI sound practices consultation (June 2026)FSBThe cross-border baseline the FSB is converging on.Proposed
SR 23-4Federal ReserveThe model provider is a third party: due diligence, contract terms, monitoring, exit.In force
FSB AI monitoring report (Oct 2025)FSBConcentration on a small number of model suppliers.Final
OCC Bulletin 2026-13OCCFor a national bank, the OCC's copy of the 2026 model-risk guidance.In force
NIST AI RMF 1.0NISTThe voluntary Govern, Map, Measure, Manage frame for everything model-risk guidance leaves out.In force

How will you know it works, before and after launch?

CODE CHECKS50%JUDGE MODEL30%HUMAN REVIEW20%
  • A golden dataset of at least 150 real cases with expected outputs, including adversarial inputs: wrong documents, unusual formats, prompts that try to change the task.
  • Code-based checks on every output: schema conformance, required fields, reconciliations against the system of record. Threshold: 99% or better before launch, every run in production.
  • Citation checks: every factual claim resolves to a passage in the governed set; unsupported claims below 2% of answers.
  • The same suite reruns on every prompt change, model version change and retrieval change; a regression blocks the release. That is what ongoing monitoring and outcomes analysis mean in model-risk terms.

Where must a person be in the loop?

  • Per-action approval by a competent reviewer for anything customer-facing or irreversible; sampled review for the rest.
  • Validation proportionate to materiality, with monitoring for drift on inputs and outputs.
  • An escalation route to a person that the customer can reach in one step.
  • Monthly review of the evals and the exception log by the accountable owner.
  • AI datasets assembled outside the governed perimeter.
  • Lineage the model asserts but the catalogue cannot reproduce.
  • Automated responses to data-subject requests without review.

What will a validator or an examiner ask?

  1. Where is this system in your inventory, what tier did you assign, and who signed it off?
  2. What counts as a model here, and what does your validation cover for the parts that are not?
  3. Show me the data lineage behind the retrieval set and the training or tuning data.
  4. What can the system do without a person, and where is that written down?
  5. How do you know it is still working: which evals run, how often, and what happened the last time one failed?
  6. What did you do about the vendor: due diligence, contract, exit plan, concentration?
  7. Walk me through one wrong output from production and what the customer, if any, saw.
  8. Who can switch it off, and has that been tested?

What it does: classifies, governs and answers questions about personal and risk data; the model prepares and may act within limits; a person approves anything that reaches a customer or cannot be undone.

Pattern: augmented call; tier 2: act with approval.

Rules it answers to: 4 documents across several jurisdictions, each linked in the brief.

How we know it works: a golden dataset, automated checks on every release, and human review at the level the tier demands.

What could go wrong and who answers: the accountable owner, the kill switch, the escalation route.

Which of the 100 largest US banks have put this use case on the record?

Every rule this brief cites, the morning it changes.

agent deployments, regulator positions and the day's six stories · in your inbox by 7 am ET · free

plus every tracker, bank and agent page update, the morning after · leave any morning