Banks are past the pilot stage with AI agents and before the platform stage. Supervisors now describe the same practice on both sides of the Atlantic — agents limited to defined use cases, behind guardrails, with a human accountable — while the rules that reach agents are scattered across model-risk, third-party, cyber, consumer-protection and data frameworks. This page assembles them into one operating model: eight control layers every agent runs on, an eight-stage lifecycle with a gate before each stage, a map of where agents are landing by bank function, a five-level autonomy ladder that says what oversight each level owes, and a dated account of how regulators are moving. Every claim links to a primary source in the regulation tracker.
What does an operating system for AI agents in a bank look like?
The term is a design goal, not a regulatory one. It means the controls an agent needs — identity, permissions, data access, model management, runtime limits, logging, oversight — are provided once as shared services and inherited by every agent, instead of being re-implemented, unevenly, inside each use case. The layers below are the minimum set. The top and bottom layers are about people and policy; the six in between are enforced in code.
LAYER 0 · WHO OWNS THE OUTCOME
Governance and accountability
Who is accountable when an AI agent acts on behalf of the bank?
LAYER 1 · WHO THE AGENT IS
Identity and entitlements
Does an AI agent need its own identity and access rights?
LAYER 2 · WHAT THE AGENT MAY DO
Action gateway
How does a bank limit what an AI agent can actually do?
LAYER 4 · WHAT THE AGENT THINKS WITH
Models and vendors
Are AI agents covered by model risk management or third-party risk management?
LAYER 5 · WHERE THE AGENT RUNS
Runtime and orchestration
What does a safe runtime for AI agents look like?
LAYER 6 · WHAT THE BANK CAN SEE
Observability, evaluation and audit
What must a bank log and monitor about its AI agents?
LAYER 7 · WHEN A PERSON STEPS IN
Human oversight and escalation
What does 'human in the loop' have to mean for an AI agent to satisfy regulators?
Read the control plane in full → — each layer's question, a quotable answer, the controls, and the documents it answers to.
How does an AI agent get from idea to production in a bank?
Through eight stages, each closed by a gate question. The sequence is the one the Treasury's December 2024 report asked for — review every use case for compliance before deployment, then re-evaluate periodically — laid over the model-risk, third-party and deployer duties that already exist. The re-assessment loop is not optional: vendor model updates, new uses, new rules and incidents all send an agent back through tiering.
| # | Stage | Gate question |
|---|---|---|
| 1 | Intake | What will the agent do, for whom, and what could go wrong? |
| 2 | Risk tiering | Which tier, which autonomy level, which rulebook? |
| 3 | Design | What is the permission envelope and where does a human sit? |
| 4 | Build and onboard | Do we know what we are buying, and can we leave? |
| 5 | Validate and test | Does it do what we approved, and what happens when attacked? |
| 6 | Approve and deploy | Who signs, and is the oversight actually configured? |
| 7 | Operate and monitor | Is it still doing what we approved? |
| 8 | Change, incident and retirement | What changed, what broke, and when do we re-assess? |
Read the lifecycle in full → — evidence per gate and the documents behind each.
Where are banks deploying AI agents today?
Everywhere the work is text, cases and exceptions — and furthest where supervisors are pushing rather than restraining. The EBA found 92% of EU banks deploying AI in 2025 and 55% already using general-purpose or agentic AI with consumers; the ECB reported in February 2026 that more than 85% of large European banks use AI, with generative and agentic tools accelerating in IT operations, legal and document analysis and front-line support; the CFPB found every top-10 US bank running a chatbot as early as 2023. The UK's 2024 survey put fully autonomous use cases at 2%. Fraud and security operations run the most autonomous agents because that is where regulators actively encourage AI.
| Function | What agents do | Level | Sources |
|---|---|---|---|
| Customer service Front office | Chat and voice agents that answer, triage, and execute routine servicing; the most widely deployed agentic use in banking. | L2 | CFPB Chatbots in Consumer Finance (issue spotlight, 2023) · EBA report: Rising application of AI in EU banking and payments (Sep 2025) |
| Onboarding and KYC Front office | Document collection, identity checks and case preparation; the front line against AI-generated identity fraud. | L2 | FIN-2024-Alert004 (Deepfake Media) · EBA report: Rising application of AI in EU banking and payments (Sep 2025) |
| Lending and underwriting Front office | Agents that assemble applications, pre-screen and explain decisions; the credit decision itself stays under consumer law and, in the EU, the high-risk regime. | L1 | ECOA / Regulation B adverse action (15 U.S.C. 1691(d); 12 CFR 1002.9) · Regulation (EU) 2024/1689 · SB 26-189 |
| Advice and wealth Front office | Research, portfolio commentary and client preparation for advisers; automated recommendations are examined for accuracy and suitability. | L1 | Division of Examinations FY2026 Priorities |
| Fraud and AML Middle office | Alert triage, investigation drafting and case narratives; the FSB's consultation includes agentic fraud detection at a large bank as a case study. | L3 | FSB AI sound practices consultation (June 2026) · Hill House oversight testimony (Jun 2026) · Division of Examinations FY2026 Priorities |
| Credit and market risk Middle office | Analysis, scenario narration and model documentation; the quantitative models remain under model risk management. | L1 | SR 26-2 · Cook: Opportunities and Risks of AI (May 2026) |
| Compliance and regulatory change Middle office | Regulatory-document analysis, obligation mapping and policy drafting — the ECB's most-cited generative use after IT operations. | L2 | Machado speech: 'Technology is neutral, governance is not' (Feb 2026) · EBA report: Rising application of AI in EU banking and payments (Sep 2025) |
| Treasury and ALM Middle office | Forecast commentary, funding memos and counterparty document review around models that stay fully validated. | L1 | SR 26-2 · BCBS 239 |
| Operations and payments Back office | Exception handling, reconciliation and payment-investigation agents inside critical operations. | L2 | BCBS Principles for Operational Resilience (2021) · FIN-2024-Alert004 (Deepfake Media) |
| Finance and reporting Back office | Report drafting, variance analysis and data-quality checks over regulated risk data. | L1 | BCBS 239 · BCBS 239 Implementation Newsletter (Jan 2026) |
| IT and engineering Back office | Coding agents, incident management and change automation — the fastest-growing generative use in European banks, with human code review now a supervisory ask. | L2 | Machado speech: 'Technology is neutral, governance is not' (Feb 2026) · DFS Frontier AI Models Industry Letter (May 2026) · Cook: Opportunities and Risks of AI (May 2026) |
| Cyber and security operations Back office | Detection, triage and response agents in the SOC — the one place supervisors are actively asking banks to use AI faster. | L3 | NIST IR 8596 (Cyber AI Profile) · ESA Statement on ICT risks from frontier AI models (JC 2026 25) · OCC Semiannual Risk Perspective, Spring 2026 |
How much autonomy should an AI agent in a bank have?
As much as the oversight around it can carry. The consistent regulatory position — the FSB's Sound Practice 10, the EU AI Act's Article 14, Colorado's human-review right, the OCC's observed ‘human-in-the-loop accountability’ — is that human oversight scales up with autonomy rather than being designed out. The ladder makes that explicit: each level names the oversight it owes.
| Level | What the agent does | Oversight it owes | Sources |
|---|---|---|---|
| L0 · Inform | Retrieves, summarises and answers. Takes no action on any system. | Accuracy testing and source lineage; a route to a human where customers are involved. | CFPB Chatbots in Consumer Finance (issue spotlight, 2023) · NIST AI 600-1 (Generative AI Profile) |
| L1 · Draft | Proposes an action, a document or code. A person executes it. | Human review before anything leaves the drafting stage — the standard supervisors set for AI-generated code. | DFS Frontier AI Models Industry Letter (May 2026) |
| L2 · Act with approval | Executes each action only after a person approves it. | Per-action approval by a competent reviewer, logged — the 'human-in-the-loop accountability' the OCC observed at banks in 2026. | OCC Semiannual Risk Perspective, Spring 2026 · Regulation (EU) 2024/1689 |
| L3 · Act within bounds | Executes autonomously inside a pre-approved envelope: allow-listed tools, value and rate limits, sampled review. | Extra oversight measures scaled to autonomy, a kill switch, continuous monitoring and human review on request. | FSB AI sound practices consultation (June 2026) · CAISI RFI on AI agent security (2026) · SB 26-189 |
| L4 · Autonomous | Closed-loop, multi-step, self-directed operation across systems. | Rare in banking: 2% of use cases were fully autonomous in the UK's 2024 survey. The FSB names autonomous multi-step action as a risk and the FSB Chair flagged frontier models' autonomy to the G20. | 2026 BoE/FCA AI survey · FSB Chair's letter to G20 (Aug 2026) |
How is the regulation of AI agents in banking moving?
In two directions at once. Enablement: adoption data, a pro-innovation federal posture, a voluntary sector framework from Treasury, and a modernised model-risk standard that deliberately carved agents out. Controls: agent-specific security work at NIST, human-oversight duties in the FSB's practices and in statute, and a run of frontier-AI cyber warnings from New York, Frankfurt, Paris and Basel in the summer of 2026. The next dated events are the FSB's final practices in October 2026, Colorado on January 1, 2027, and the EU AI Act's high-risk obligations on December 2, 2027.
How mature is agentic AI in banking?
Stage three, on the supervisors' own description: governed agents in defined use cases with guardrails and measured oversight. The move to stage four — controls as shared platform services — is what turns a collection of governed agents into an operating system, and it is where the engineering investment is going. Stage five, multi-agent operations at high autonomy, is where the standards are still being drafted.
STAGE 1
Experiments
Pilots in individual teams. No inventory, no tiering, shared credentials, vendor terms unread.
- Nobody can list every AI-enabled feature in production
- Agents run on human or service-account credentials
STAGE 2
Copilots at scale
Level 0–1 assistants deployed widely: summarisation, drafting, coding. An inventory exists; governance is policy, not enforcement.
- Inventory and owners exist
- Human review of outputs is expected but not evidenced
STAGE 3
Governed agents
Level 2–3 agents in defined use cases with guardrails, approval steps and measured human oversight. This is the practice supervisors describe at large banks in 2026.
- Tiering decides the control set
- Approvals and overrides are logged and reported
- Vendor AI under third-party risk management
STAGE 4
Agent operating system
Identity, gateway, data access, runtime and observability are shared platform services; policy is enforced in code across every agent rather than re-implemented per use case.
- One identity and entitlement model for agents
- One gateway, one trace store, one evaluation pipeline
- New agents inherit controls by default
STAGE 5
Multi-agent operations
Agents coordinate across functions at level 3–4 autonomy. The controls exist only in outline: NIST's multi-agent overlay is its least mature, and the FSB names autonomous multi-step action as a risk.
- Declared, traceable hand-offs between agents
- Budgets and kill switches at the workflow level, not only per agent
Which documents does this model cite?
Every document above has its own page with the official link, a summary and what changed. See the tracker for all 19 authorities.
What is an operating system for AI agents in a bank?
The shared set of controls every AI agent runs on, rather than controls re-implemented per use case: governance and accountability, agent identity and entitlements, an action gateway that limits what agents can do, entitlement-aware data access with lineage, model and vendor management, a sandboxed and budgeted runtime, observability and audit, and human oversight with escalation. The term is a design goal — no regulator prescribes it — but each layer maps to documents supervisors already cite.
Are AI agents covered by bank model risk management?
Not in the United States since April 17, 2026. The revised interagency model risk guidance (SR 26-2, OCC Bulletin 2026-13, FDIC FIL-15-2026) says generative and agentic AI models are outside its scope and must be managed through broader risk-management and governance programs. The traditional models an agent calls — credit, fraud, liquidity — remain in scope, and vendor-supplied agents fall under third-party guidance (SR 23-4). The agencies have promised a request for information on AI and model risk.
What autonomy level do banks actually run today?
Mostly levels 1 and 2 — drafting and acting with approval. The OCC reported in May 2026 that bank use of generative and agentic AI is primarily productivity and customer-experience tools with guardrails and human-in-the-loop accountability. The UK's 2024 survey found 2% of use cases fully autonomous. Level 3, acting within a pre-approved envelope, is emerging in fraud and security operations, where supervisors are most encouraging.
What do regulators say about agentic AI specifically?
Four things, consistently. It is not yet inside model risk guidance (Fed, OCC, FDIC). Human oversight must scale up with autonomy (FSB Sound Practice 10, EU AI Act Article 14, Colorado). Agent-specific security — prompt injection, memory poisoning, AI-generated code — is an active workstream (NIST CAISI, DFS, the ESAs). And frontier models' autonomy is now a financial-stability topic (FSB Chair's August 2026 letter).
Which documents should an AI agent program cite?
For US banks: SR 26-2 / Bulletin 2026-13 for the models inside the agent and the carve-out around it; SR 23-4 for vendor AI; NIST AI 600-1 for generative-AI risks; the CAISI agent-security RFI for agent-specific threats; Treasury's FS AI RMF as the assessment template. Internationally: the FSB's 12 sound practices, the Basel third-party principles, BCBS 239 for data, the EU AI Act's Articles 12, 14 and 26, and the ESAs' July 2026 statement under DORA.
Agents move fast. So do their regulators.
6 curated AI stories for banking executives · Every morning · Free
Subscribe to BankingNewsAI →