All posts
Methodology · · 7 min read

What AI should not do in a loan review

AI can do most of the analytical work. What it should never do is absorb accountability for the judgments an examiner will ask about.

Most writing about AI in loan review is about what it can do. That list is genuinely long now, and growing: reading a credit file, spreading financials, normalizing cash flow, proposing a sample against a methodology, drafting narrative, flagging exceptions and missing documents.

The more useful list is the short one on the other side. It is worth being precise about what kind of list that is. The 2020 Interagency Guidance on Credit Risk Review Systems is clear that credit risk review should be carried out by qualified personnel with appropriate experience and training, and that it should be an independent assessment rather than a repackaging of information produced elsewhere. What it does not do is legislate which parts of the analytical work a piece of software may perform.

So what follows is Klerum's governance position rather than a regulatory recitation: a deliberately conservative structure we think is defensible in an examination, and which we have built the product around. Four responsibilities stay with qualified people.

1. Approving the final risk rating

Start with how rating decisions actually work, because it is more distributed than most AI commentary assumes. Lending staff generally hold primary responsibility for assigning timely risk ratings. Credit risk review then validates those ratings independently and, where warranted, adjusts them. Where independent review concludes credit quality is worse than management believes, the lower rating typically prevails unless management produces information that satisfies the reviewer or an arbiter.

Depending on the institution, the reviewer may assign a review rating, recommend a downgrade, establish the controlling classification, refer a disagreement to a committee, or require management to update the system of record. The common thread is not that a human types the grade. It is that an authorized person evaluates the evidence and owns the outcome.

Against that, a well-governed AI can do a great deal: assemble what the decision needs (the trend, the covenant position, the global cash flow, the comparison to the officer's original grade), flag where the file and the assigned grade appear inconsistent, and calculate or recommend a rating under the institution's approved methodology. Banks have used scorecards and rules engines in rating processes for years, and nothing in the guidance forbids that.

The control that matters is approval authority, not authorship. And there is a practical risk worth naming: a default AI rating with passive override rights tends to produce human confirmation rather than genuine human judgment. Under deadline, an approval step that requires no engagement with the evidence is not really a control. That is an argument about how the workflow should be designed, not a claim that regulators have prohibited the software from proposing a grade.

2. Approving the review scope

Scope is a governance decision, and the governance is distributed. Credit review policy addresses the frequency, scope, and depth of reviews. Institution personnel, usually credit review management, develop the review plan. The board or an appropriate board committee typically approves the scope annually and when significant changes occur. Individual reviewers make engagement-level adjustments, but they do not own the institution's overall coverage methodology.

By materiality here we mean two specific things: the coverage thresholds that determine how much of the portfolio gets reviewed, and the significance threshold that determines when an observation is raised as a finding. Both are institutional settings, not per-file reviewer preferences.

AI is excellent at executing an approved scope: applying weighting rules across thousands of loans consistently, which is exactly where manual sampling tends to drift. It should propose against the approved methodology, show its work, and let credit review management adjust while the board or committee approves the overall scope. What it should not become is the reason the scope is what it is. “The system selected it” is not an answer to “why was this portfolio under-covered?”

A model can tell you what the file says. It cannot be accountable for what the file concludes.

3. Material findings and conclusions on credit administration

The conclusions that matter most, whether risk-rating practices are effective, whether credit administration is adequate, and whether the credit information supporting the allowance process is reliable, are professional opinions. They carry weight with a board and an examiner precisely because a qualified person formed them and is answerable for them.

That framing is deliberate. Credit risk review supplies management with accurate and timely credit-quality information for determining the allowance, and it assesses the quality of risk identification and grading. Concluding on the soundness of the institution's entire allowance methodology is a broader exercise that typically involves finance, model risk management, validation, internal audit, and external audit depending on the bank. Loan review contributes to it rather than owning it.

On drafting: AI may propose language, and it may propose a potential finding or conclusion for the reviewer to test. The control objective is not that the model must stay silent until a human has decided. It is that the reviewer independently assesses the evidence against source documents and institutional standards rather than approving the model's default. An AI-drafted paragraph asserting that credit administration is satisfactory has the grammar of a conclusion; it acquires the standing of one only when a qualified person has genuinely evaluated it.

4. Signing the completed workpaper

Someone must be answerable, and accountability is the one thing that cannot be delegated to a model or a vendor. As a practical control, every material finding and final conclusion should carry a named reviewer and a recorded approval, and that record should be structural rather than a workflow convention.

This also answers a question raised by SR 26-2 placing generative AI outside the scope of model risk guidance. Plenty of governance remains available: third-party risk management, information security, operational risk, credit review policy, change management, software testing, and independent testing designed for generative systems. But because SR 26-2 does not prescribe a validation framework for these tools, documented human acceptance of material conclusions becomes an especially important control.

Where the line sits

TaskAIAuthorized person
Read the credit file and extract figuresYes, with citationsVerifies at source
Normalize and spread cash flowYesConfirms treatment of adjustments
Propose a sample against the approved methodologyYesCredit review management adjusts; board or committee approves overall scope
Flag file and grade inconsistenciesYesInvestigates
Recommend or calculate a risk ratingYes, if governed and explainableReviews the supporting evidence
Draft finding and memo narrativeYes, including proposed conclusionsIndependently assesses, then edits and owns the text
Approve the institution's final ratingNoDecides and documents
Set coverage and finding-significance thresholdsNoInstitution sets; board or committee approves
Conclude on credit administrationNoDecides
Sign the workpaperNoSigns

What to ask about an autonomous review pitch

Vendors occasionally pitch autonomous review: files in, finished report out. It makes a compelling demo, and it is worth slowing down to ask what the resulting file actually establishes.

A claim of end-to-end autonomous loan review should prompt specific questions. Who is independent of the lending function in this process? Who holds approval authority for a rating, and what evidence did they see before exercising it? What happens when the review conclusion and management's assessment disagree, and who arbitrates? What does the institution hand an examiner who asks how a conclusion was reached? Sophisticated buyers generally reach their own conclusion once those questions are on the table.

The underlying point is about product design rather than anyone's competence. The value of a loan review is not the document; it is that an independent, qualified party evaluated the credit and is prepared to stand behind what they found. A system that produces the document while removing the evaluation has produced a very fast summary of a credit file, which is a different and less valuable thing.

The right ambition is narrower and more useful: take the search, transcription, and assembly out of the work so that the judgment, the part only a person can supply, gets more of the engagement rather than less.

This piece describes Klerum's governance position and general practice under the 2020 Interagency Guidance on Credit Risk Review Systems. It is not legal or compliance advice, and supervisory expectations vary by institution and charter.

Sources

See Klerum on your own portfolio.

A 30-minute demo, no deck. Upload a sample trial and watch the AI propose a review population against your methodology.