AI agents in banking · The build map · Brief 02

Fair lending & discrimination: a design brief for banks.

Last updated Sep 17, 2026 · standard answers for this use case · adjust them in the map

Fair-lending work is analytical and evidentiary: the outputs are tests, comparisons and documentation that will be read by an examiner or a court. That makes it a workflow with an evaluator, not an agent. Models help by drafting the analysis plan, explaining a model's features in plain language and assembling the file; the statistical tests run in code, and a person owns the conclusion.

WorkflowPrompt chainingEvaluator and optimizerTier 2: act with approvalautonomy level 2

Fair lending & discrimination: The model prepares and may act within limits; a person approves anything that reaches a customer or cannot be undone. Several model calls orchestrated by your code, in a sequence or a graph you designed. Material but recoverable. The model may prepare and, within limits, act, with a person approving anything that leaves the bank or touches a customer.

PatternWorkflow. Your code lays out the steps, the cost of error is real, and you need to observe every step.Prompt chainingEach step's output is the next step's input, with a programmatic check between them.Evaluator and optimizerOne model produces, another checks against explicit criteria, and the loop repeats until the check passes or a person is called.KnowledgeRetrieval over the governed document set, with a citation on every answer and a refusal when nothing relevant is found. Tool calls to the system of record for anything live. Never retrieval for a balance, a case status or a limit. Documents to reason with, systems to check against: the workflow decides which question goes where before the model answers.Delivery routeProvider API: one third party to diligence; residency and retention terms are yours to negotiate.

Which steps belong to a person, the model and a system?

A PERSONTHE MODELA SYSTEM1
Define protected classes and the comparison design
2
Run disparate-impact tests
3
Explain features and interactions
4
Draft the memo
5
Conclude and remediate
#StepOwnerNote
1Define protected classes and the comparison designA personCounsel and compliance set the frame.
2Run disparate-impact testsA systemDeterministic statistics on the decision data.
3Explain features and interactionsThe modelPlain-language explanations of what the tests found, cited to the outputs.
4Draft the memoThe modelFrom the results; nothing asserted beyond them.
5Conclude and remediateA personSigned, with the actions taken.

Which control layers carry the weight?

Governance and accountabilityCOREIdentity and entitlementsCOREAction gatewayCOREData and knowledgeCOREModels and vendorsLIGHTRuntime and orchestrationSTANDARDObservability, evaluation and auditSTANDARDHuman oversight and escalationSTANDARD

Each layer is described, with its controls and documents, on the control plane page.

Which rules and guidance does this design answer to?

DocumentAuthorityWhy it applies hereStatus
Regulation B final rule on disparate impact (April 2026)CFPBThe current US position on disparate impact under Regulation B.In force
ECOA / Regulation B adverse action (15 U.S.C. 1691(d); 12 CFR 1002.9)CFPBAdverse-action reasons are also the audit trail for discrimination testing.In force
Joint Statement on Automated Systems (CFPB, DOJ, EEOC, FTC)CFPBFour agencies on automated systems and discrimination law.Final
Regulation (EU) 2024/1689EU AI ActBias examination is a data-governance duty for high-risk systems (Article 10).In force
SR 26-2Federal ReserveUS model risk management as revised in April 2026.In force
SR 23-4Federal ReserveUS third-party risk management, including the model provider.In force
SB 26-189Colorado AI ActColorado: notice, explanation and human review for consequential automated decisions from Jan 1, 2027.Final
FSB AI monitoring report (Oct 2025)FSBConcentration on a small number of model suppliers.Final
OCC Bulletin 2026-13OCCFor a national bank, the OCC's copy of the 2026 model-risk guidance.In force
NIST AI RMF 1.0NISTThe voluntary Govern, Map, Measure, Manage frame for everything model-risk guidance leaves out.In force

How will you know it works, before and after launch?

CODE CHECKS30%JUDGE MODEL50%HUMAN REVIEW20%
  • A golden dataset of at least 150 real cases with expected outputs, including adversarial inputs: wrong documents, unusual formats, prompts that try to change the task.
  • A judge model scoring against a written rubric (accuracy, completeness, tone, citation present), calibrated against a human-scored sample every month. Threshold set from the human sample, not guessed.
  • Citation checks: every factual claim resolves to a passage in the governed set; unsupported claims below 2% of answers.
  • Escalation evals: the cases that must reach a person do, on a held-out set, with precision and recall both reported.
  • Trace evals per step, not only end to end: which step fails, how often, at what cost, so a prompt or model change can be judged step by step.
  • The same suite reruns on every prompt change, model version change and retrieval change; a regression blocks the release. That is what ongoing monitoring and outcomes analysis mean in model-risk terms.

Where must a person be in the loop?

  • Per-action approval by a competent reviewer for anything customer-facing or irreversible; sampled review for the rest.
  • Validation proportionate to materiality, with monitoring for drift on inputs and outputs.
  • An escalation route to a person that the customer can reach in one step.
  • Monthly review of the evals and the exception log by the accountable owner.
  • Adverse-action explainability is tested as a launch gate: if the system cannot produce specific principal reasons, it does not decide.
  • Using a model to run or interpret the statistics without the code path to reproduce them.
  • Drift in the bias tests after a model change with no rerun.
  • Treating an explainability tool's output as the legal explanation.

What will a validator or an examiner ask?

  1. Where is this system in your inventory, what tier did you assign, and who signed it off?
  2. What counts as a model here, and what does your validation cover for the parts that are not?
  3. Show me the data lineage behind the retrieval set and the training or tuning data.
  4. What can the system do without a person, and where is that written down?
  5. How do you know it is still working: which evals run, how often, and what happened the last time one failed?
  6. What did you do about the vendor: due diligence, contract, exit plan, concentration?
  7. Walk me through one wrong output from production and what the customer, if any, saw.
  8. Who can switch it off, and has that been tested?
  9. Produce the adverse-action reasons for this declined applicant from the decision record.

What it does: tests models and decisions for disparate treatment and impact, and documents the results; the model prepares and may act within limits; a person approves anything that reaches a customer or cannot be undone.

Pattern: workflow (prompt chaining, evaluator and optimizer); tier 2: act with approval.

Rules it answers to: 4 documents across US, each linked in the brief.

How we know it works: a golden dataset, automated checks on every release, and human review at the level the tier demands.

What could go wrong and who answers: the accountable owner, the kill switch, the escalation route.

Every rule this brief cites, the morning it changes.

agent deployments, regulator positions and the day's six stories · in your inbox by 7 am ET · free

plus every tracker, bank and agent page update, the morning after · leave any morning