Credit is the use case with the clearest rulebook and the least room for autonomy. The scoring model itself is a model in the model-risk sense and stays under validation; what the newer patterns add is everything around it: assembling the file from documents and systems, checking it for completeness, drafting the decision memo and producing the reasons a declined applicant is owed. The design that survives an exam keeps the decision with a person or a validated model, and uses the language model to prepare and explain, never to decide on its own.
Credit scoring & underwriting: The model prepares, drafts and checks. A person decides and acts, and the record shows who. Several model calls orchestrated by your code, in a sequence or a graph you designed. High stakes, hard to undo, customer-facing. The model prepares, drafts and checks; a person decides and acts, and the record shows it.
Which steps belong to a person, the model and a system?
| # | Step | Owner | Note |
|---|---|---|---|
| 1 | Collect and classify documents | The model | Extraction against a schema, with a confidence per field. |
| 2 | Verify income, identity and existing exposure | A system | Tool calls to the core, the bureau and the KYC store; never retrieval. |
| 3 | Score and price | A system | The validated scoring model, outside the language model entirely. |
| 4 | Draft the decision memo | The model | From the file and the score; every number traceable to its source. |
| 5 | Decide | A person | The underwriter, with the memo; the record shows the reviewer. |
| 6 | Adverse-action notice | A system | Reasons generated from the decision record and checked against the required list. |
Which control layers carry the weight?
Each layer is described, with its controls and documents, on the control plane page.
- Governance and accountability: core for this design.
- Identity and entitlements: core for this design.
- Action gateway: core for this design.
- Data and knowledge: core for this design.
- Models and vendors: core for this design.
- Observability, evaluation and audit: core for this design.
- Human oversight and escalation: core for this design.
Which rules and guidance does this design answer to?
| Document | Authority | Why it applies here | Status |
|---|---|---|---|
| ECOA / Regulation B adverse action (15 U.S.C. 1691(d); 12 CFR 1002.9) | CFPB | Specific principal reasons for any adverse credit action, whatever the model. | In force |
| FCRA adverse action and credit-score disclosures (15 U.S.C. 1681m, 1681g(f)) | CFPB | Key factors behind a credit score must be disclosed. | In force |
| SR 26-2 | Federal Reserve | A credit model is a model: validation and monitoring scaled to materiality. | In force |
| Regulation (EU) 2024/1689 | EU AI Act | Credit scoring of natural persons is high-risk in the EU (Annex III 5(b)). | In force |
| EBA Guidelines on loan origination and monitoring (EBA/GL/2020/06) | EBA | EU rulebook for automated creditworthiness models; staff must be able to override. | In force |
| CFPB Chatbots in Consumer Finance (issue spotlight, 2023) | CFPB | Customer-facing AI must not block access to a person or give wrong answers about rights. | Final |
| SR 23-4 | Federal Reserve | US third-party risk management, including the model provider. | In force |
| SB 26-189 | Colorado AI Act | Colorado: notice, explanation and human review for consequential automated decisions from Jan 1, 2027. | Final |
| BCBS Third-Party Risk Principles (Dec 2025) | Basel Committee | Nth-party supply chains and concentration on cloud providers. | In force |
| BCBS ICT Risk Management Report (June 2026) | Basel Committee | How supervisors look at ICT and cloud dependencies. | Final |
| OCC Bulletin 2026-13 | OCC | For a national bank, the OCC's copy of the 2026 model-risk guidance. | In force |
| NIST AI RMF 1.0 | NIST | The voluntary Govern, Map, Measure, Manage frame for everything model-risk guidance leaves out. | In force |
How will you know it works, before and after launch?
- A golden dataset of at least 300 real cases with expected outputs, including adversarial inputs: wrong documents, unusual formats, prompts that try to change the task.
- A judge model scoring against a written rubric (accuracy, completeness, tone, citation present), calibrated against a human-scored sample every month. Threshold set from the human sample, not guessed.
- Citation checks: every factual claim resolves to a passage in the governed set; unsupported claims below 2% of answers.
- Escalation evals: the cases that must reach a person do, on a held-out set, with precision and recall both reported.
- Trace evals per step, not only end to end: which step fails, how often, at what cost, so a prompt or model change can be judged step by step.
- The same suite reruns on every prompt change, model version change and retrieval change; a regression blocks the release. That is what ongoing monitoring and outcomes analysis mean in model-risk terms.
Where must a person be in the loop?
- A named accountable executive signs off the use case before build, and the inventory records the tier.
- A person reviews every output before it reaches a customer or a system of record; the review is logged with the reviewer's identity.
- Independent validation before launch and outcomes analysis after it, on the model-risk schedule for a material model.
- Adverse outcomes to a customer carry the specific reasons the law requires, produced from the decision record, not reconstructed afterwards.
- A kill switch and a rollback path, tested before launch.
- The customer is told they are dealing with an AI system and can reach a person in one step.
- Adverse-action explainability is tested as a launch gate: if the system cannot produce specific principal reasons, it does not decide.
- Letting the language model influence the score, which makes the scoring model unvalidatable.
- Reconstructing adverse-action reasons after the fact instead of from the decision record.
- Forgetting that a fraud-driven decline is still an adverse action.
What will a validator or an examiner ask?
- Where is this system in your inventory, what tier did you assign, and who signed it off?
- What counts as a model here, and what does your validation cover for the parts that are not?
- Show me the data lineage behind the retrieval set and the training or tuning data.
- What can the system do without a person, and where is that written down?
- How do you know it is still working: which evals run, how often, and what happened the last time one failed?
- What did you do about the vendor: due diligence, contract, exit plan, concentration?
- Walk me through one wrong output from production and what the customer, if any, saw.
- Who can switch it off, and has that been tested?
- Produce the adverse-action reasons for this declined applicant from the decision record.
- Show me how a customer reaches a person, and how long it took the last ten who tried.
What it does: assembles and checks a credit application, and explains a decision; the model prepares, drafts and checks. a person decides and acts, and the record shows who.
Pattern: workflow (parallelization, evaluator and optimizer); tier 1: a person decides.
Rules it answers to: 4 documents across US, each linked in the brief.
How we know it works: a golden dataset, automated checks on every release, and human review at the level the tier demands.
What could go wrong and who answers: the accountable owner, the kill switch, the escalation route.
Which of the 100 largest US banks have put this use case on the record?
Every rule this brief cites, the morning it changes.
agent deployments, regulator positions and the day's six stories · in your inbox by 7 am ET · free
plus every tracker, bank and agent page update, the morning after · leave any morning