
A 4-question check you can run against the source page during the demo.

Every claims AI evaluation runs into the same question: how do you know the answer is true? This sets out a four-question check you can run against the source page, the two forms an AI error actually takes in a claim file, and what a citation has to do before it counts as evidence.
Buyers ask the same question before every claims AI demo. How do you know the answer is true?
That is the correct question. It is also answerable, but not by a vendor assurance. It is answerable by a procedure the buyer can run during the demo.
This post sets out that procedure. It explains what an AI error looks like in a claim file. It explains why a citation can make an error harder to catch. And it sets out what a source link must do before it counts as evidence.
The word covers two failures that behave differently in practice.
Research from Stanford Law School measured both in legal tools. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools tested three retrieval-based products. The tools from LexisNexis and Thomson Reuters "each hallucinate between 17% and 33% of the time."
Those are retrieval-based tools with real document databases behind them. Retrieval reduces the error rate. It does not remove it.
The system states something the record does not contain. A date, a diagnosis, a provider, a figure.
This form is the easier one. Anyone who knows the file spots it.
The paper calls this misgrounded: a response where "key factual propositions are cited but the source does not support the claim." The citation is real. The link between the citation and the claim is not.
Its definition of hallucination covers both forms: "A response is considered hallucinated if it is either incorrect or misgrounded." A false statement counts. So does a false claim that a source supports a statement.
The second form is the dangerous one in a claim file. The page number is real. The document is real. The link opens. Only the connection between them is wrong.
This is the part buyers get backwards. A cited answer feels more trustworthy, so it gets checked less.
The Stanford authors put it directly. These errors "are potentially more dangerous than fabricating a case outright, because they are subtler and more difficult to spot."
They also describe what real verification costs. Checking requires users "to click through to cited references, read and understand the relevant sources." It then requires them to "assess their authority, and compare them to the propositions the model seeks to support."
That is the whole job. The value of a source link is not that it exists. It is that it makes those four steps fast enough to actually perform.
Run these in order on any answer that would change a decision.
Open it. Confirm the page is in the production you are working from, at the number given.
A link that opens a viewer without landing on a specific page fails here. So does a citation to a document the file does not contain.
Read the passage, not the page heading. Compare the words on the page with the words in the answer.
Three failures recur. The page supports a weaker version of the claim. The page supports the claim for a different date or body part. Or the page is the source of a quotation the answer has paraphrased into something stronger.
A claim about a diagnosis should cite a clinical note or an imaging report. A claim about a charge should cite a bill or a ledger.
An intake form recording what the claimant said is evidence of what the claimant said. It is not evidence that the thing is true. Answers that blur this are the most common quiet failure in claim file summaries.
A correct citation can still produce a misleading answer. The file may contain a contradicting document that the answer ignores. In general liability claims, where several parties file their own documents, this is routine.
Ask the system directly for contradicting material. An answer that can only support its own position is a search result, not an analysis.
Take an answer of a kind that appears in almost every auto liability claim file.
"The claimant had no treatment between 14 March and 2 September." (source: p. 212)
Question one. Page 212 opens, and it is a physical therapy discharge note dated 14 March. It exists.
Question two. The page says therapy was discontinued on 14 March. It does not say anything about September. The answer has combined two facts and cited only one of them.
That is the misgrounded form. Nothing on page 212 is false. The proposition it was cited for is not the proposition it supports.
Question three. A discharge note is the right document type for "treatment ended." It is the wrong document type for "no treatment occurred anywhere for six months." That claim needs the absence of records across every provider, which is a claim about the whole file.
Question four. The billing summary at page 340 shows two chiropractic charges in June. The gap is real but shorter, and the answer as written would not survive the demand response. The answer took one minute to check. The version in the file now reads differently, and its citations hold up.
Vendor demos run on prepared files. Three tests survive that.
Bring a file you have worked. Ask five questions you can answer yourself, and check every source link.
The point is not whether the system gets them right. It is whether you can tell, and how long telling takes.
Ask for a document the file does not contain. Ask for a diagnosis that was never made.
A system that answers confidently anyway has told you what happens on the questions you cannot check. A system that says the file does not contain it has told you something better.
Generation speed is the least interesting number in the demo. Verification time is the one that changes your day.
Measure from reading the answer to confirming the cited page supports it. If that takes several minutes per answer, the review has moved rather than shrunk. Platforms where every response is traceable to its source are aiming at this number specifically.
Not all citations are equal. A usable one meets five conditions.
Structured claim file review for bodily injury is judged on these mechanics rather than on the fluency of the summary. A well-written answer with weak sourcing is worse than a plain one with strong sourcing.
Verification is not only an internal quality question. It is a documented expectation.
The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers was adopted on 4 December 2023. It sets out state regulators' expectations for insurers using AI systems. It asks for a written AI systems programme, governance, testing and validation, and oversight of third parties acting for the insurer.
It also tells insurers what may be requested later. The bulletin "advises insurers of documentation that a state Department of Insurance may request during an investigation or examination."
Federal guidance points the same way. NIST AI 600-1, the Generative AI Profile published on 26 July 2024, is a companion to the AI Risk Management Framework. It is voluntary, and it is written to help organisations build trustworthiness into how AI systems are evaluated and used.
Neither document asks whether the vendor promised accuracy. Both ask what the insurer did to check.
Verifying every sentence of every answer defeats the purpose. Verifying nothing defeats the file.
A workable split has three tiers.
Log what was checked. The log is what turns individual diligence into a programme, and it is what an examiner or opposing counsel will ask to see.
Source checking catches wrong and misgrounded answers. It does not catch three other things.
It does not catch a complete file that is missing a provider nobody requested. It does not catch a question nobody thought to ask. And it does not catch a judgement call that was documented correctly and made badly.
Those stay with the adjuster, which is where they belong. Verification tells you the facts in front of you are real. Deciding what they mean is still the job.