# Building a system that triages alerts, assembles cases and drafts investigations for AML and KYC teams inside a bank

Source: https://www.bankingnewsai.com/agentic-banking/build/aml-kyc
Last updated: Sep 17, 2026

AML is where supervisors have been most encouraging and where the volume argument is strongest: thousands of alerts, most of them false positives, each needing a file. The winning pattern is routing plus evaluator: classify the alert, gather the evidence from the case system and the transaction store in parallel, draft the narrative, then check it against the evidence before an investigator sees it. The investigator still disposes of the alert and signs the report; the model removes hours of assembly.

## The brief (default answers)

| Choice | Answer |
| --- | --- |
| Who is affected | Staff only |
| Reversibility | With cost |
| Stakes | Material |
| Knowledge | Both |
| Verifiability | By judgment |
| Steps | Partly |
| Jurisdiction | United States |
| Delivery route | Through a cloud platform |

- **Pattern:** Workflow — Several model calls orchestrated by your code, in a sequence or a graph you designed.
- **Tier:** 2, Tier 2: act with approval — Material but recoverable. The model may prepare and, within limits, act, with a person approving anything that leaves the bank or touches a customer.
- **Workflow shapes:** Routing, Evaluator and optimizer
- **Human involvement:** The model prepares and may act within limits; a person approves anything that reaches a customer or cannot be undone.
- **Knowledge:** Retrieval over the governed document set, with a citation on every answer and a refusal when nothing relevant is found.; Tool calls to the system of record for anything live. Never retrieval for a balance, a case status or a limit.; Documents to reason with, systems to check against: the workflow decides which question goes where before the model answers.

## Decomposition

| Step | Owner | Note |
| --- | --- | --- |
| Classify the alert | model | Routing to a typology-specific path. |
| Gather evidence | system | Tool calls to the case system, transactions and KYC store, in parallel. |
| Draft the investigation narrative | model | Every statement cited to an evidence item. |
| Check the draft against the evidence | model | An evaluator pass with explicit criteria; failures go back or to a person. |
| Dispose and file | human | The investigator decides; the SAR is theirs. |

## Control layer weights (0–4)

| Layer | Weight |
| --- | --- |
| governance | 1 |
| identity | 1 |
| actions | 0.9 |
| data | 1 |
| models | 1 |
| runtime | 0.8 |
| observability | 0.7 |
| oversight | 0.6 |

## Documents that apply

- [2018 Joint Statement on BSA/AML Innovation](https://www.bankingnewsai.com/ai-regulation/documents/fincen-joint-statement-innovation-2018) — Agencies encourage innovative approaches, including AI, in BSA/AML programmes.
- [2026 AML/CFT Program Proposed Rule](https://www.bankingnewsai.com/ai-regulation/documents/fincen-aml-cft-program-nprm-2026) — 'Effective use of artificial intelligence' counted in an institution's favour, if finalised.
- [SR 26-2](https://www.bankingnewsai.com/ai-regulation/documents/fed-sr-26-2) — Replaces the 2021 BSA/AML model-risk statement; monitoring models are models.
- [FIN-2024-Alert004 (Deepfake Media)](https://www.bankingnewsai.com/ai-regulation/documents/fincen-alert-2024-deepfake-media) — Deepfake documents at onboarding are the threat KYC models now face.
- [SR 23-4](https://www.bankingnewsai.com/ai-regulation/documents/fed-sr-23-4) — US third-party risk management, including the model provider.
- [SB 26-189](https://www.bankingnewsai.com/ai-regulation/documents/co-sb26-189) — Colorado: notice, explanation and human review for consequential automated decisions from Jan 1, 2027.
- [BCBS Third-Party Risk Principles (Dec 2025)](https://www.bankingnewsai.com/ai-regulation/documents/bcbs-third-party-risk-principles-2025) — Nth-party supply chains and concentration on cloud providers.
- [BCBS ICT Risk Management Report (June 2026)](https://www.bankingnewsai.com/ai-regulation/documents/bcbs-ict-risk-management-range-of-practices-2026) — How supervisors look at ICT and cloud dependencies.
- [OCC Bulletin 2026-13](https://www.bankingnewsai.com/ai-regulation/documents/occ-bulletin-2026-13) — For a national bank, the OCC's copy of the 2026 model-risk guidance.
- [NIST AI RMF 1.0](https://www.bankingnewsai.com/ai-regulation/documents/nist-ai-100-1) — The voluntary Govern, Map, Measure, Manage frame for everything model-risk guidance leaves out.

## Evals

- A golden dataset of at least 150 real cases with expected outputs, including adversarial inputs: wrong documents, unusual formats, prompts that try to change the task.
- A judge model scoring against a written rubric (accuracy, completeness, tone, citation present), calibrated against a human-scored sample every month. Threshold set from the human sample, not guessed.
- Citation checks: every factual claim resolves to a passage in the governed set; unsupported claims below 2% of answers.
- Trace evals per step, not only end to end: which step fails, how often, at what cost, so a prompt or model change can be judged step by step.
- The same suite reruns on every prompt change, model version change and retrieval change; a regression blocks the release. That is what ongoing monitoring and outcomes analysis mean in model-risk terms.

## Human gates

- Per-action approval by a competent reviewer for anything customer-facing or irreversible; sampled review for the rest.
- Validation proportionate to materiality, with monitoring for drift on inputs and outputs.
- An escalation route to a person that the customer can reach in one step.
- Monthly review of the evals and the exception log by the accountable owner.

## What an examiner will ask

- Where is this system in your inventory, what tier did you assign, and who signed it off?
- What counts as a model here, and what does your validation cover for the parts that are not?
- Show me the data lineage behind the retrieval set and the training or tuning data.
- What can the system do without a person, and where is that written down?
- How do you know it is still working: which evals run, how often, and what happened the last time one failed?
- What did you do about the vendor: due diligence, contract, exit plan, concentration?
- Walk me through one wrong output from production and what the customer, if any, saw.
- Who can switch it off, and has that been tested?

## What the board should hear

- What it does: triages alerts, assembles cases and drafts investigations for AML and KYC teams; the model prepares and may act within limits; a person approves anything that reaches a customer or cannot be undone.
- Pattern: workflow (routing, evaluator and optimizer); tier 2: act with approval.
- Rules it answers to: 4 documents across US, each linked in the brief.
- How we know it works: a golden dataset, automated checks on every release, and human review at the level the tier demands.
- What could go wrong and who answers: the accountable owner, the kill switch, the escalation route.

## Pitfalls

- Letting the model close alerts on its own; the 2026 proposal rewards effective AI use, not unsupervised disposal.
- Retrieval over stale customer data instead of a live tool call.
- No record of what the model read, which makes the file indefensible.

Change any answer on the interactive map: https://www.bankingnewsai.com/agentic-banking/build/aml-kyc

---

Canonical page: https://www.bankingnewsai.com/agentic-banking/build/aml-kyc
Part of [BankingNewsAI](https://www.bankingnewsai.com/) — a free daily brief on AI in banking, an AI regulation tracker (19 authorities, 166 documents) and AI-strategy profiles of the 100 largest US banks. Markdown versions of every reference page: append `.md` to the page URL; index at https://www.bankingnewsai.com/llms.txt.
