AI agents in banking · A working model

An operating system for
AI agents in your bank.

Last updated Sep 6, 2026 · 41 primary-source documents cited · Updated as regulators move

Banks are past the pilot stage with AI agents and before the platform stage. Supervisors now describe the same practice on both sides of the Atlantic — agents limited to defined use cases, behind guardrails, with a human accountable — while the rules that reach agents are scattered across model-risk, third-party, cyber, consumer-protection and data frameworks. This page assembles them into one operating model: eight control layers every agent runs on, an eight-stage lifecycle with a gate before each stage, a map of where agents are landing by bank function, a five-level autonomy ladder that says what oversight each level owes, and a dated account of how regulators are moving. Every claim links to a primary source in the regulation tracker.

What does an operating system for AI agents in a bank look like?

The term is a design goal, not a regulatory one. It means the controls an agent needs — identity, permissions, data access, model management, runtime limits, logging, oversight — are provided once as shared services and inherited by every agent, instead of being re-implemented, unevenly, inside each use case. The layers below are the minimum set. The top and bottom layers are about people and policy; the six in between are enforced in code.

The eight control layers an AI agent in a bank runs onLAYERWHAT IT DECIDESANSWERS TO0Governance and accountabilityWho owns the outcomeFSB AI sound practices consultation (June 2026)Machado speech: 'Technology is neutral, governance is not' (Feb 2026)OCC Bulletin 2026-131Identity and entitlementsWho the agent isCAISI RFI on AI agent security (2026)NIST COSAiS control overlays23 NYCRR Part 5002Action gatewayWhat the agent may doOCC Semiannual Risk Perspective, Spring 2026FSB AI sound practices consultation (June 2026)Regulation (EU) 2024/16893Data and knowledgeWhat the agent may knowBCBS 239BCBS 239 Implementation Newsletter (Jan 2026)NIST AI 600-1 (Generative AI Profile)4Models and vendorsWhat the agent thinks withSR 26-2SR 23-4BCBS Third-Party Risk Principles (Dec 2025)5Runtime and orchestrationWhere the agent runsCAISI RFI on AI agent security (2026)FSB AI sound practices consultation (June 2026)NIST COSAiS control overlays6Observability, evaluation and auditWhat the bank can seeRegulation (EU) 2024/1689DFS Frontier AI Models Industry Letter (May 2026)FSB AI sound practices consultation (June 2026)7Human oversight and escalationWhen a person steps inSB 26-189CPPA ADMT, risk-assessment and cybersecurity-audit regulationsRegulation (EU) 2024/1689POLICYENFORCED IN CODEPEOPLE
Figure 1 · The control plane. Governance sets the policy; layers 1–6 enforce it in code; human oversight closes the loop. Right column: the primary documents each layer answers to.

LAYER 0 · WHO OWNS THE OUTCOME

Governance and accountability

Who is accountable when an AI agent acts on behalf of the bank?

LAYER 1 · WHO THE AGENT IS

Identity and entitlements

Does an AI agent need its own identity and access rights?

LAYER 2 · WHAT THE AGENT MAY DO

Action gateway

How does a bank limit what an AI agent can actually do?

LAYER 3 · WHAT THE AGENT MAY KNOW

Data and knowledge

What data controls do AI agents need?

LAYER 4 · WHAT THE AGENT THINKS WITH

Models and vendors

Are AI agents covered by model risk management or third-party risk management?

LAYER 5 · WHERE THE AGENT RUNS

Runtime and orchestration

What does a safe runtime for AI agents look like?

LAYER 6 · WHAT THE BANK CAN SEE

Observability, evaluation and audit

What must a bank log and monitor about its AI agents?

LAYER 7 · WHEN A PERSON STEPS IN

Human oversight and escalation

What does 'human in the loop' have to mean for an AI agent to satisfy regulators?

Read the control plane in full → — each layer's question, a quotable answer, the controls, and the documents it answers to.

How does an AI agent get from idea to production in a bank?

Through eight stages, each closed by a gate question. The sequence is the one the Treasury's December 2024 report asked for — review every use case for compliance before deployment, then re-evaluate periodically — laid over the model-risk, third-party and deployer duties that already exist. The re-assessment loop is not optional: vendor model updates, new uses, new rules and incidents all send an agent back through tiering.

The eight-stage lifecycle of an AI agent, with a gate before every stage01Intake02Risk tiering03Design04Build andonboard05Validate andtest06Approve anddeploy07Operate andmonitor08Change,incident andretirementRE-ASSESS · VENDOR UPDATE · NEW USE · NEW RULE · INCIDENT◇ GATE QUESTION■ SIGN-OFF: ACCOUNTABLE OWNER + SECOND LINE
Figure 2 · The lifecycle. Each diamond is a gate question the stage must answer before the next begins; the return arrow is the re-assessment loop that vendor updates, new uses, new rules and incidents all trigger.
#StageGate question
1IntakeWhat will the agent do, for whom, and what could go wrong?
2Risk tieringWhich tier, which autonomy level, which rulebook?
3DesignWhat is the permission envelope and where does a human sit?
4Build and onboardDo we know what we are buying, and can we leave?
5Validate and testDoes it do what we approved, and what happens when attacked?
6Approve and deployWho signs, and is the oversight actually configured?
7Operate and monitorIs it still doing what we approved?
8Change, incident and retirementWhat changed, what broke, and when do we re-assess?

Read the lifecycle in full → — evidence per gate and the documents behind each.

Where are banks deploying AI agents today?

Everywhere the work is text, cases and exceptions — and furthest where supervisors are pushing rather than restraining. The EBA found 92% of EU banks deploying AI in 2025 and 55% already using general-purpose or agentic AI with consumers; the ECB reported in February 2026 that more than 85% of large European banks use AI, with generative and agentic tools accelerating in IT operations, legal and document analysis and front-line support; the CFPB found every top-10 US bank running a chatbot as early as 2023. The UK's 2024 survey put fully autonomous use cases at 2%. Fraud and security operations run the most autonomous agents because that is where regulators actively encourage AI.

Where AI agents are being deployed across a bank, by function and autonomy levelFRONT OFFICECustomer servicechat and voice servicingL2Onboarding and KYCdocument collection, identity checksL2Lending and underwritingapplication assembly, decision explanationL1Advice and wealthresearch and adviser preparationL1MIDDLE OFFICEFraud and AMLalert triage, investigation draftingL3Credit and market riskanalysis and model documentationL1Compliance and regulatory changeobligation mapping, policy draftingL2Treasury and ALMforecast commentary, funding memosL1BACK OFFICEOperations and paymentsexceptions, reconciliation, investigationsL2Finance and reportingreport drafting, data-quality checksL1IT and engineeringcoding agents, incident managementL2Cyber and security operationsdetection, triage, responseL3
Figure 3 · Where agents are landing. Dots show the autonomy level that is observed or defensible today (0 inform · 1 draft · 2 act with approval · 3 act within bounds · 4 autonomous). Fraud and security operations lead because supervisors actively encourage AI there; lending stays at level 1 because the decision is governed by consumer law.
FunctionWhat agents doLevelSources
Customer service
Front office
Chat and voice agents that answer, triage, and execute routine servicing; the most widely deployed agentic use in banking.L2CFPB Chatbots in Consumer Finance (issue spotlight, 2023) · EBA report: Rising application of AI in EU banking and payments (Sep 2025)
Onboarding and KYC
Front office
Document collection, identity checks and case preparation; the front line against AI-generated identity fraud.L2FIN-2024-Alert004 (Deepfake Media) · EBA report: Rising application of AI in EU banking and payments (Sep 2025)
Lending and underwriting
Front office
Agents that assemble applications, pre-screen and explain decisions; the credit decision itself stays under consumer law and, in the EU, the high-risk regime.L1ECOA / Regulation B adverse action (15 U.S.C. 1691(d); 12 CFR 1002.9) · Regulation (EU) 2024/1689 · SB 26-189
Advice and wealth
Front office
Research, portfolio commentary and client preparation for advisers; automated recommendations are examined for accuracy and suitability.L1Division of Examinations FY2026 Priorities
Fraud and AML
Middle office
Alert triage, investigation drafting and case narratives; the FSB's consultation includes agentic fraud detection at a large bank as a case study.L3FSB AI sound practices consultation (June 2026) · Hill House oversight testimony (Jun 2026) · Division of Examinations FY2026 Priorities
Credit and market risk
Middle office
Analysis, scenario narration and model documentation; the quantitative models remain under model risk management.L1SR 26-2 · Cook: Opportunities and Risks of AI (May 2026)
Compliance and regulatory change
Middle office
Regulatory-document analysis, obligation mapping and policy drafting — the ECB's most-cited generative use after IT operations.L2Machado speech: 'Technology is neutral, governance is not' (Feb 2026) · EBA report: Rising application of AI in EU banking and payments (Sep 2025)
Treasury and ALM
Middle office
Forecast commentary, funding memos and counterparty document review around models that stay fully validated.L1SR 26-2 · BCBS 239
Operations and payments
Back office
Exception handling, reconciliation and payment-investigation agents inside critical operations.L2BCBS Principles for Operational Resilience (2021) · FIN-2024-Alert004 (Deepfake Media)
Finance and reporting
Back office
Report drafting, variance analysis and data-quality checks over regulated risk data.L1BCBS 239 · BCBS 239 Implementation Newsletter (Jan 2026)
IT and engineering
Back office
Coding agents, incident management and change automation — the fastest-growing generative use in European banks, with human code review now a supervisory ask.L2Machado speech: 'Technology is neutral, governance is not' (Feb 2026) · DFS Frontier AI Models Industry Letter (May 2026) · Cook: Opportunities and Risks of AI (May 2026)
Cyber and security operations
Back office
Detection, triage and response agents in the SOC — the one place supervisors are actively asking banks to use AI faster.L3NIST IR 8596 (Cyber AI Profile) · ESA Statement on ICT risks from frontier AI models (JC 2026 25) · OCC Semiannual Risk Perspective, Spring 2026

How much autonomy should an AI agent in a bank have?

As much as the oversight around it can carry. The consistent regulatory position — the FSB's Sound Practice 10, the EU AI Act's Article 14, Colorado's human-review right, the OCC's observed ‘human-in-the-loop accountability’ — is that human oversight scales up with autonomy rather than being designed out. The ladder makes that explicit: each level names the oversight it owes.

Five levels of AI agent autonomy and the oversight each one requiresL0InformAccuracy tests, lineage,a route to a humanL1DraftA person reviews beforeanything is executedL2Act with approvalPer-action approval by acompetent reviewer, loggedL3Act within boundsEnvelope, sampled review,kill switch, monitoringL4AutonomousRare in banking; risksnamed by FSB and NISTWHERE BANKS ARE · OCC, MAY 20262% OF USE CASES · UK 2024OVERSIGHT REQUIRED ↓MORE
Figure 4 · The autonomy ladder. Oversight scales up with autonomy, not down: that is the FSB's Sound Practice 10, the EU AI Act's Article 14 and Colorado's human-review right in one picture. Shaded: where supervisors say banks are in 2026.
LevelWhat the agent doesOversight it owesSources
L0 · InformRetrieves, summarises and answers. Takes no action on any system.Accuracy testing and source lineage; a route to a human where customers are involved.CFPB Chatbots in Consumer Finance (issue spotlight, 2023) · NIST AI 600-1 (Generative AI Profile)
L1 · DraftProposes an action, a document or code. A person executes it.Human review before anything leaves the drafting stage — the standard supervisors set for AI-generated code.DFS Frontier AI Models Industry Letter (May 2026)
L2 · Act with approvalExecutes each action only after a person approves it.Per-action approval by a competent reviewer, logged — the 'human-in-the-loop accountability' the OCC observed at banks in 2026.OCC Semiannual Risk Perspective, Spring 2026 · Regulation (EU) 2024/1689
L3 · Act within boundsExecutes autonomously inside a pre-approved envelope: allow-listed tools, value and rate limits, sampled review.Extra oversight measures scaled to autonomy, a kill switch, continuous monitoring and human review on request.FSB AI sound practices consultation (June 2026) · CAISI RFI on AI agent security (2026) · SB 26-189
L4 · AutonomousClosed-loop, multi-step, self-directed operation across systems.Rare in banking: 2% of use cases were fully autonomous in the UK's 2024 survey. The FSB names autonomous multi-step action as a risk and the FSB Chair flagged frontier models' autonomy to the G20.2026 BoE/FCA AI survey · FSB Chair's letter to G20 (Aug 2026)

How is the regulation of AI agents in banking moving?

In two directions at once. Enablement: adoption data, a pro-innovation federal posture, a voluntary sector framework from Treasury, and a modernised model-risk standard that deliberately carved agents out. Controls: agent-specific security work at NIST, human-oversight duties in the FSB's practices and in statute, and a run of frontier-AI cyber warnings from New York, Frankfurt, Paris and Basel in the summer of 2026. The next dated events are the FSB's final practices in October 2026, Colorado on January 1, 2027, and the EU AI Act's high-risk obligations on December 2, 2027.

Timeline of regulatory and market events shaping AI agents in banking, 2023 to 2027ENABLEMENTCONTROLS AND WARNINGS2023NIST AI RMF 1.0JAN 2023CFPB on bank chatbotsJUN 2023SR 23-4 third-party guidanceJUN 20232024Treasury AI cyber and fraud reportMAR 2024EU AI Act publishedJUL 2024NIST AI 600-1JUL 2024FinCEN deepfake alertNOV 2024FSB: six AI vulnerabilitiesNOV 2024Treasury RFI reportDEC 20242025EBA: 92% of EU banks deploy AISEP 2025Basel third-party principlesDEC 2025FSOC AI Working GroupDEC 20252026NIST CAISI agent-security RFIJAN 2026Treasury FS AI RMFFEB 2026ECB: 'governance is not neutral'FEB 2026SR 26-2 / Bulletin 2026-13APR 2026OCC Risk PerspectiveMAY 2026DFS frontier-AI letterMAY 2026UK survey asks about agentsJUN 2026FSB 12 sound practicesJUN 2026HM Treasury adoption planJUL 2026ESAs on frontier AI under DORAJUL 2026FSB Chair to the G20AUG 2026FSB final sound practicesOCT 2026 · EXPECTED2027Colorado ADMT Act in forceJAN 2027 · EXPECTEDEU AI Act high-risk obligationsDEC 2027 · EXPECTED
Figure 5 · How things are moving. Left: enablement — adoption data and pro-innovation policy. Right: controls — guidance, warnings and statutes. Hollow markers are scheduled or expected. Every event links to its primary source in the tracker.

Read what each regulator has actually said about agents →

How mature is agentic AI in banking?

Stage three, on the supervisors' own description: governed agents in defined use cases with guardrails and measured oversight. The move to stage four — controls as shared platform services — is what turns a collection of governed agents into an operating system, and it is where the engineering investment is going. Stage five, multi-agent operations at high autonomy, is where the standards are still being drafted.

Five maturity stages from AI experiments to multi-agent operations01Experiments02Copilots atscale03Governed agentsLARGE BANKS · 202604Agent operatingsystemWHERE THIS IS GOING05Multi-agentoperationsWHERE THIS IS GOINGCONTROLS MOVE FROM POLICY → PER-AGENT → PLATFORM
Figure 6 · Maturity. Stage 3 is what supervisors describe at large banks in 2026; stages 4 and 5 are where the platform work is heading, and where the standards (NIST's agent overlays, the FSB's final practices) are still being written.

STAGE 1

Experiments

Pilots in individual teams. No inventory, no tiering, shared credentials, vendor terms unread.

  • Nobody can list every AI-enabled feature in production
  • Agents run on human or service-account credentials

STAGE 2

Copilots at scale

Level 0–1 assistants deployed widely: summarisation, drafting, coding. An inventory exists; governance is policy, not enforcement.

  • Inventory and owners exist
  • Human review of outputs is expected but not evidenced

STAGE 3

Governed agents

Level 2–3 agents in defined use cases with guardrails, approval steps and measured human oversight. This is the practice supervisors describe at large banks in 2026.

  • Tiering decides the control set
  • Approvals and overrides are logged and reported
  • Vendor AI under third-party risk management

STAGE 4

Agent operating system

Identity, gateway, data access, runtime and observability are shared platform services; policy is enforced in code across every agent rather than re-implemented per use case.

  • One identity and entitlement model for agents
  • One gateway, one trace store, one evaluation pipeline
  • New agents inherit controls by default

STAGE 5

Multi-agent operations

Agents coordinate across functions at level 3–4 autonomy. The controls exist only in outline: NIST's multi-agent overlay is its least mature, and the FSB names autonomous multi-step action as a risk.

  • Declared, traceable hand-offs between agents
  • Budgets and kill switches at the workflow level, not only per agent

Which documents does this model cite?

AuthorityDocuments
FSBFSB Chair's letter to G20 (Aug 2026) · FSB AI sound practices consultation (June 2026) · FSB AI monitoring report (Oct 2025) · FSB AI financial stability report (Nov 2024)
ECBMachado speech: 'Technology is neutral, governance is not' (Feb 2026)
OCCOCC Semiannual Risk Perspective, Spring 2026 · OCC Bulletin 2026-13
NISTCAISI RFI on AI agent security (2026) · NIST IR 8596 (Cyber AI Profile) · NIST COSAiS control overlays · NIST AI 100-2e2025 (Adversarial ML) · NIST AI 600-1 (Generative AI Profile) · NIST AI RMF 1.0
U.S. TreasuryTreasury FS AI RMF and AI Lexicon (Feb 2026) · FSOC 2025 Annual Report · Treasury AI in Financial Services report (Dec 2024) · Treasury AI cybersecurity risks report (Mar 2024)
NY DFSDFS Frontier AI Models Industry Letter (May 2026) · 23 NYCRR Part 500
EU AI ActRegulation (EU) 2024/1689
Colorado AI ActSB 26-189
CFPBCFPB Chatbots in Consumer Finance (issue spotlight, 2023) · ECOA / Regulation B adverse action (15 U.S.C. 1691(d); 12 CFR 1002.9)
Basel CommitteeBCBS 239 Implementation Newsletter (Jan 2026) · BCBS Third-Party Risk Principles (Dec 2025) · BCBS Principles for Operational Resilience (2021) · BCBS 239
Federal ReserveCook: Opportunities and Risks of AI (May 2026) · Bowman: AI in the Financial System (May 2026) · SR 26-2 · SR 23-4
EBAESA Statement on ICT risks from frontier AI models (JC 2026 25) · EBA report: Rising application of AI in EU banking and payments (Sep 2025)
SECDivision of Examinations FY2026 Priorities
California CPPACPPA ADMT, risk-assessment and cybersecurity-audit regulations
UK (BoE / PRA / FCA)HM Treasury Financial Services AI Adoption Plan (Jul 2026) · 2026 BoE/FCA AI survey · FCA FS25/5 · PRA SS1/23
FinCENFIN-2024-Alert004 (Deepfake Media)
FDICHill House oversight testimony (Jun 2026)

Every document above has its own page with the official link, a summary and what changed. See the tracker for all 19 authorities.

What is an operating system for AI agents in a bank?

The shared set of controls every AI agent runs on, rather than controls re-implemented per use case: governance and accountability, agent identity and entitlements, an action gateway that limits what agents can do, entitlement-aware data access with lineage, model and vendor management, a sandboxed and budgeted runtime, observability and audit, and human oversight with escalation. The term is a design goal — no regulator prescribes it — but each layer maps to documents supervisors already cite.

Are AI agents covered by bank model risk management?

Not in the United States since April 17, 2026. The revised interagency model risk guidance (SR 26-2, OCC Bulletin 2026-13, FDIC FIL-15-2026) says generative and agentic AI models are outside its scope and must be managed through broader risk-management and governance programs. The traditional models an agent calls — credit, fraud, liquidity — remain in scope, and vendor-supplied agents fall under third-party guidance (SR 23-4). The agencies have promised a request for information on AI and model risk.

What autonomy level do banks actually run today?

Mostly levels 1 and 2 — drafting and acting with approval. The OCC reported in May 2026 that bank use of generative and agentic AI is primarily productivity and customer-experience tools with guardrails and human-in-the-loop accountability. The UK's 2024 survey found 2% of use cases fully autonomous. Level 3, acting within a pre-approved envelope, is emerging in fraud and security operations, where supervisors are most encouraging.

What do regulators say about agentic AI specifically?

Four things, consistently. It is not yet inside model risk guidance (Fed, OCC, FDIC). Human oversight must scale up with autonomy (FSB Sound Practice 10, EU AI Act Article 14, Colorado). Agent-specific security — prompt injection, memory poisoning, AI-generated code — is an active workstream (NIST CAISI, DFS, the ESAs). And frontier models' autonomy is now a financial-stability topic (FSB Chair's August 2026 letter).

Which documents should an AI agent program cite?

For US banks: SR 26-2 / Bulletin 2026-13 for the models inside the agent and the carve-out around it; SR 23-4 for vendor AI; NIST AI 600-1 for generative-AI risks; the CAISI agent-security RFI for agent-specific threats; Treasury's FS AI RMF as the assessment template. Internationally: the FSB's 12 sound practices, the Basel third-party principles, BCBS 239 for data, the EU AI Act's Articles 12, 14 and 26, and the ESAs' July 2026 statement under DORA.

Agents move fast. So do their regulators.

6 curated AI stories for banking executives · Every morning · Free

Subscribe to BankingNewsAI →