In markets the risks are disclosure and supervision: claims about AI must match what the systems do, communications must be supervised as communications, and anything that touches order flow is a model under validation. The safe patterns are augmented calls and chained workflows for research synthesis, surveillance triage and draft communications, with a person approving anything that goes to a client and nothing autonomous near execution.
Trading & capital markets: The model prepares and may act within limits; a person approves anything that reaches a customer or cannot be undone. Several model calls orchestrated by your code, in a sequence or a graph you designed. Material but recoverable. The model may prepare and, within limits, act, with a person approving anything that leaves the bank or touches a customer.
Which steps belong to a person, the model and a system?
| # | Step | Owner | Note |
|---|---|---|---|
| 1 | Synthesise research | The model | From licensed sources and internal notes, with citations. |
| 2 | Triage surveillance alerts | The model | Summaries and proposed dispositions; the compliance officer decides. |
| 3 | Draft client communications | The model | Reviewed and supervised like any other communication. |
| 4 | Execute | A system | Validated algorithms with their own controls; no language model in the path. |
| 5 | Describe AI use to clients and regulators | A person | Accurately; the AI-washing cases were about the gap between claim and capability. |
Which control layers carry the weight?
Each layer is described, with its controls and documents, on the control plane page.
- Governance and accountability: core for this design.
- Identity and entitlements: core for this design.
- Action gateway: core for this design.
- Data and knowledge: core for this design.
Which rules and guidance does this design answer to?
| Document | Authority | Why it applies here | Status |
|---|---|---|---|
| Division of Examinations FY2026 Priorities | SEC | Examiners test whether AI claims are accurate and AI use is supervised. | In force |
| CFTC Staff Advisory 24-17 on AI | CFTC | CFTC staff on AI use by registrants: existing rules apply. | In force |
| Delphia / Global Predictions AI-washing settlements | SEC | The first 'AI washing' settlements: claims must match capability. | Final |
| SR 26-2 | Federal Reserve | Trading models validated as models, scaled to materiality. | In force |
| SR 23-4 | Federal Reserve | US third-party risk management, including the model provider. | In force |
| SB 26-189 | Colorado AI Act | Colorado: notice, explanation and human review for consequential automated decisions from Jan 1, 2027. | Final |
| FSB AI monitoring report (Oct 2025) | FSB | Concentration on a small number of model suppliers. | Final |
| OCC Bulletin 2026-13 | OCC | For a national bank, the OCC's copy of the 2026 model-risk guidance. | In force |
| NIST AI RMF 1.0 | NIST | The voluntary Govern, Map, Measure, Manage frame for everything model-risk guidance leaves out. | In force |
How will you know it works, before and after launch?
- A golden dataset of at least 150 real cases with expected outputs, including adversarial inputs: wrong documents, unusual formats, prompts that try to change the task.
- A judge model scoring against a written rubric (accuracy, completeness, tone, citation present), calibrated against a human-scored sample every month. Threshold set from the human sample, not guessed.
- Citation checks: every factual claim resolves to a passage in the governed set; unsupported claims below 2% of answers.
- Escalation evals: the cases that must reach a person do, on a held-out set, with precision and recall both reported.
- Trace evals per step, not only end to end: which step fails, how often, at what cost, so a prompt or model change can be judged step by step.
- The same suite reruns on every prompt change, model version change and retrieval change; a regression blocks the release. That is what ongoing monitoring and outcomes analysis mean in model-risk terms.
Where must a person be in the loop?
- Per-action approval by a competent reviewer for anything customer-facing or irreversible; sampled review for the rest.
- Validation proportionate to materiality, with monitoring for drift on inputs and outputs.
- An escalation route to a person that the customer can reach in one step.
- Monthly review of the evals and the exception log by the accountable owner.
- Marketing that overstates what AI does in the process.
- Generated communications outside the supervision and retention system.
- A language model with a route, however indirect, to order entry.
What will a validator or an examiner ask?
- Where is this system in your inventory, what tier did you assign, and who signed it off?
- What counts as a model here, and what does your validation cover for the parts that are not?
- Show me the data lineage behind the retrieval set and the training or tuning data.
- What can the system do without a person, and where is that written down?
- How do you know it is still working: which evals run, how often, and what happened the last time one failed?
- What did you do about the vendor: due diligence, contract, exit plan, concentration?
- Walk me through one wrong output from production and what the customer, if any, saw.
- Who can switch it off, and has that been tested?
What it does: supports research, surveillance and client communication in markets businesses; the model prepares and may act within limits; a person approves anything that reaches a customer or cannot be undone.
Pattern: workflow (prompt chaining, evaluator and optimizer); tier 2: act with approval.
Rules it answers to: 4 documents across US, each linked in the brief.
How we know it works: a golden dataset, automated checks on every release, and human review at the level the tier demands.
What could go wrong and who answers: the accountable owner, the kill switch, the escalation route.
Which of the 100 largest US banks have put this use case on the record?
Every rule this brief cites, the morning it changes.
agent deployments, regulator positions and the day's six stories · in your inbox by 7 am ET · free
plus every tracker, bank and agent page update, the morning after · leave any morning