Security is the one area where regulators are urging speed: the ECB, New York's DFS and the OCC have all said AI-enabled attackers are faster, and expect AI-enabled defence in return. The pattern is routing and parallelization on live telemetry: enrich an alert from many sources at once, classify, propose a response, and let a person approve any containment action that could disrupt the business. AI-generated code and configuration go through human review before deployment, which DFS has said in so many words.
Cybersecurity: The model prepares and may act within limits; a person approves anything that reaches a customer or cannot be undone. Several model calls orchestrated by your code, in a sequence or a graph you designed. Material but recoverable. The model may prepare and, within limits, act, with a person approving anything that leaves the bank or touches a customer.
Which steps belong to a person, the model and a system?
| # | Step | Owner | Note |
|---|---|---|---|
| 1 | Enrich the alert | A system | Parallel tool calls to logs, identity, endpoint and threat intelligence. |
| 2 | Classify and summarise | The model | Severity, likely technique, affected assets, with sources. |
| 3 | Propose containment | The model | A ranked set of actions with expected impact. |
| 4 | Approve and execute | A person | Containment that could disrupt service needs a person; low-impact actions by rule. |
| 5 | Draft the incident record | The model | From the trace; reviewed before it becomes the record. |
Which control layers carry the weight?
Each layer is described, with its controls and documents, on the control plane page.
- Governance and accountability: core for this design.
- Identity and entitlements: core for this design.
- Action gateway: core for this design.
- Models and vendors: core for this design.
Which rules and guidance does this design answer to?
| Document | Authority | Why it applies here | Status |
|---|---|---|---|
| ECB 'Dear CEO' letter on AI-enabled cybersecurity threats (SSM-2026-0301) | ECB | AI models that find and exploit vulnerabilities; action plans due Oct 31, 2026. | In force |
| DFS Frontier AI Models Industry Letter (May 2026) | NY DFS | Human review of AI-generated code before deployment; shorter patch cycles. | In force |
| 23 NYCRR Part 500 | NY DFS | The New York cybersecurity regulation AI guidance hangs on. | In force |
| ESA Statement on ICT risks from frontier AI models (JC 2026 25) | EBA | EU supervisors on ICT risk from frontier models. | In force |
| Treasury AI cybersecurity risks report (Mar 2024) | U.S. Treasury | Treasury's map of AI-specific cyber risk for the sector. | Final |
| SR 26-2 | Federal Reserve | US model risk management as revised in April 2026. | In force |
| SR 23-4 | Federal Reserve | US third-party risk management, including the model provider. | In force |
| SB 26-189 | Colorado AI Act | Colorado: notice, explanation and human review for consequential automated decisions from Jan 1, 2027. | Final |
| BCBS Third-Party Risk Principles (Dec 2025) | Basel Committee | Nth-party supply chains and concentration on cloud providers. | In force |
| BCBS ICT Risk Management Report (June 2026) | Basel Committee | How supervisors look at ICT and cloud dependencies. | Final |
| OCC Bulletin 2026-13 | OCC | For a national bank, the OCC's copy of the 2026 model-risk guidance. | In force |
| NIST AI RMF 1.0 | NIST | The voluntary Govern, Map, Measure, Manage frame for everything model-risk guidance leaves out. | In force |
How will you know it works, before and after launch?
- A golden dataset of at least 150 real cases with expected outputs, including adversarial inputs: wrong documents, unusual formats, prompts that try to change the task.
- Code-based checks on every output: schema conformance, required fields, reconciliations against the system of record. Threshold: 99% or better before launch, every run in production.
- Trace evals per step, not only end to end: which step fails, how often, at what cost, so a prompt or model change can be judged step by step.
- The same suite reruns on every prompt change, model version change and retrieval change; a regression blocks the release. That is what ongoing monitoring and outcomes analysis mean in model-risk terms.
Where must a person be in the loop?
- Per-action approval by a competent reviewer for anything customer-facing or irreversible; sampled review for the rest.
- Validation proportionate to materiality, with monitoring for drift on inputs and outputs.
- An escalation route to a person that the customer can reach in one step.
- Monthly review of the evals and the exception log by the accountable owner.
- Automated containment with no envelope, taking down a payment system.
- AI-generated detections or patches deployed without review.
- Telemetry fed to a model outside the bank's data boundary.
What will a validator or an examiner ask?
- Where is this system in your inventory, what tier did you assign, and who signed it off?
- What counts as a model here, and what does your validation cover for the parts that are not?
- Show me the data lineage behind the retrieval set and the training or tuning data.
- What can the system do without a person, and where is that written down?
- How do you know it is still working: which evals run, how often, and what happened the last time one failed?
- What did you do about the vendor: due diligence, contract, exit plan, concentration?
- Walk me through one wrong output from production and what the customer, if any, saw.
- Who can switch it off, and has that been tested?
What it does: supports detection, triage and response in the security operations centre; the model prepares and may act within limits; a person approves anything that reaches a customer or cannot be undone.
Pattern: workflow (routing); tier 2: act with approval.
Rules it answers to: 4 documents across US, each linked in the brief.
How we know it works: a golden dataset, automated checks on every release, and human review at the level the tier demands.
What could go wrong and who answers: the accountable owner, the kill switch, the escalation route.
Which of the 100 largest US banks have put this use case on the record?
Every rule this brief cites, the morning it changes.
agent deployments, regulator positions and the day's six stories · in your inbox by 7 am ET · free
plus every tracker, bank and agent page update, the morning after · leave any morning