All posts
Methodology · · 4 min read

Citations, not accuracy percentages, are the right bar for AI in a regulated review

A 97% accuracy claim tells you nothing you can act on. A page reference tells your reviewer where to look. Only one of those survives an examination.

Every AI vendor selling into credit eventually produces a number. Ninety-four percent. Ninety-seven. Ninety-nine on a good slide. It is offered as reassurance, and it is close to meaningless in a loan review context, not because it is dishonest, but because of what it cannot tell you.

What an accuracy percentage leaves out

It was measured on documents that are not yours. Your portfolio has its own composition of borrower types, document quality, and entity structures. A figure derived from a different corpus does not transfer, and no vendor can honestly claim it does.

It averages across things you care about unequally. Extracting a borrower's name and extracting the interest expense line from a compiled statement are not equivalent tasks, and a blended score hides which one is failing. Errors in a review file are not fungible: some are cosmetic, some change a DSCR.

It says nothing about failure mode. A system that is wrong loudly, flagging uncertainty, is far safer than one that is wrong quietly with a clean-looking output. The percentage is identical in both cases.

You cannot audit it. This is the decisive objection. You cannot re-run the vendor's test, and you cannot monitor the metric on your own engagements. A control you cannot test is not a control.

An accuracy figure asks you to trust the vendor. A citation asks you to check the document. Only one of those is a control.

The question that replaces it

Reframe from “how often is the AI right?” to “when it is wrong, how quickly and cheaply does my reviewer find out?”

That reframing matters because it is the one your process can act on. You cannot make a model more accurate. You can absolutely make verification fast enough that it happens on every material figure, and once verification is routine, the underlying error rate stops being the thing standing between you and a defensible file.

This is the same logic loan review already applies to the institution it reviews. You do not ask a bank to certify that its risk ratings are 97% accurate. You sample, you trace to source, you document what you found. Applying a weaker standard to your own AI tooling than to your client's credit function would be a strange position to defend.

What a real citation looks like

The word gets used loosely, so it is worth being specific. A usable citation in loan review is:

  • Document-level and page-level. “2024 tax return” is not a citation. “2024 Form 1120S, page 3” is.
  • One click from where the reviewer is working. If the reviewer must leave the workpaper, open a file browser, and search, verification will be skipped under deadline. That is a design failure, not a discipline failure.
  • Anchored to the specific text or figure, highlighted in place, so the reviewer confirms rather than re-reads.
  • Retained in the file, so the trace survives after the engagement closes and the reviewer moves on.
Accuracy claimSource citation
Measured onVendor's corpusYour document, every time
Reproducible by youNoYes, in seconds
Per-claim or aggregateAggregatePer claim
Survives in the fileNoYes
What an examiner can testNothingAny figure in the workpaper

How to test it in a demo

Bring a real credit file, ideally one with a mediocre scan and a borrower with more than one entity. Then do three things.

Pick the least convenient figure on the screen and ask where it came from. Time how long it takes to land on the source. Then find something the system got wrong or flagged, and ask what the reviewer sees at that moment. A vendor whose product is built around verification will welcome all three. A vendor whose product is built around the percentage will steer you back to the slide.

Where the time actually goes

Verification is not the expensive part of a review. Searching is. Reading three hundred pages to locate twelve figures is what consumes the engagement; confirming twelve figures that are already located, cited, and on screen takes minutes. Cut the search cost and you can afford to verify more thoroughly than a manual process ever did, which is why a well-implemented AI review raises the evidentiary quality of the file rather than thinning it.

That is also the answer to the sceptical partner in a review firm: this is not a shortcut around professional judgment. It is a way to spend more of the engagement exercising it.

See Klerum on your own portfolio.

A 30-minute demo, no deck. Upload a sample trial and watch the AI propose a review population against your methodology.