Blog
September 27, 2026

Can You Trust a Claims AI Answer? How to Check It Against the Source Page

by
Andrej Evtimov

A 4-question check you can run against the source page during the demo.

Every claims AI evaluation runs into the same question: how do you know the answer is true? This sets out a four-question check you can run against the source page, the two forms an AI error actually takes in a claim file, and what a citation has to do before it counts as evidence.

‍

The objection is the right one

‍

Buyers ask the same question before every claims AI demo. How do you know the answer is true?

‍

That is the correct question. It is also answerable, but not by a vendor assurance. It is answerable by a procedure the buyer can run during the demo.

‍

This post sets out that procedure. It explains what an AI error looks like in a claim file. It explains why a citation can make an error harder to catch. And it sets out what a source link must do before it counts as evidence.

‍

What "hallucination" means, in two different forms

‍

The word covers two failures that behave differently in practice.

‍

Research from Stanford Law School measured both in legal tools. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools tested three retrieval-based products. The tools from LexisNexis and Thomson Reuters "each hallucinate between 17% and 33% of the time."

‍

Those are retrieval-based tools with real document databases behind them. Retrieval reduces the error rate. It does not remove it.

‍

Form one: the answer is simply wrong

‍

The system states something the record does not contain. A date, a diagnosis, a provider, a figure.

‍

This form is the easier one. Anyone who knows the file spots it.

‍

Form two: the answer cites a real page that does not support it

‍

The paper calls this misgrounded: a response where "key factual propositions are cited but the source does not support the claim." The citation is real. The link between the citation and the claim is not.

‍

Its definition of hallucination covers both forms: "A response is considered hallucinated if it is either incorrect or misgrounded." A false statement counts. So does a false claim that a source supports a statement.

‍

The second form is the dangerous one in a claim file. The page number is real. The document is real. The link opens. Only the connection between them is wrong.

‍

A citation makes an error harder to catch, not easier

‍

This is the part buyers get backwards. A cited answer feels more trustworthy, so it gets checked less.

‍

The Stanford authors put it directly. These errors "are potentially more dangerous than fabricating a case outright, because they are subtler and more difficult to spot."

‍

They also describe what real verification costs. Checking requires users "to click through to cited references, read and understand the relevant sources." It then requires them to "assess their authority, and compare them to the propositions the model seeks to support."

‍

That is the whole job. The value of a source link is not that it exists. It is that it makes those four steps fast enough to actually perform.

‍

The check: four questions against the source page

‍

Run these in order on any answer that would change a decision.

‍

Does the cited page exist in this file?

‍

Open it. Confirm the page is in the production you are working from, at the number given.

‍

A link that opens a viewer without landing on a specific page fails here. So does a citation to a document the file does not contain.

‍

Does the page say what the answer says?

‍

Read the passage, not the page heading. Compare the words on the page with the words in the answer.

‍

Three failures recur. The page supports a weaker version of the claim. The page supports the claim for a different date or body part. Or the page is the source of a quotation the answer has paraphrased into something stronger.

‍

Is this the right kind of document to prove that?

‍

A claim about a diagnosis should cite a clinical note or an imaging report. A claim about a charge should cite a bill or a ledger.

‍

An intake form recording what the claimant said is evidence of what the claimant said. It is not evidence that the thing is true. Answers that blur this are the most common quiet failure in claim file summaries.

‍

Does anything else in the file contradict it?

‍

A correct citation can still produce a misleading answer. The file may contain a contradicting document that the answer ignores. In general liability claims, where several parties file their own documents, this is routine.

‍

Ask the system directly for contradicting material. An answer that can only support its own position is a search result, not an analysis.

‍

A worked example

‍

Take an answer of a kind that appears in almost every auto liability claim file.

‍

"The claimant had no treatment between 14 March and 2 September." (source: p. 212)

‍

Question one. Page 212 opens, and it is a physical therapy discharge note dated 14 March. It exists.

‍

Question two. The page says therapy was discontinued on 14 March. It does not say anything about September. The answer has combined two facts and cited only one of them.

‍

That is the misgrounded form. Nothing on page 212 is false. The proposition it was cited for is not the proposition it supports.

‍

Question three. A discharge note is the right document type for "treatment ended." It is the wrong document type for "no treatment occurred anywhere for six months." That claim needs the absence of records across every provider, which is a claim about the whole file.

‍

Question four. The billing summary at page 340 shows two chiropractic charges in June. The gap is real but shorter, and the answer as written would not survive the demand response. The answer took one minute to check. The version in the file now reads differently, and its citations hold up.

‍

Three tests that survive a prepared demo

‍

Vendor demos run on prepared files. Three tests survive that.

‍

Ask for something you already know

‍

Bring a file you have worked. Ask five questions you can answer yourself, and check every source link.

‍

The point is not whether the system gets them right. It is whether you can tell, and how long telling takes.

‍

Ask for something that should not exist

‍

Ask for a document the file does not contain. Ask for a diagnosis that was never made.

‍

A system that answers confidently anyway has told you what happens on the questions you cannot check. A system that says the file does not contain it has told you something better.

‍

Time the verification, not the generation

‍

Generation speed is the least interesting number in the demo. Verification time is the one that changes your day.

‍

Measure from reading the answer to confirming the cited page supports it. If that takes several minutes per answer, the review has moved rather than shrunk. Platforms where every response is traceable to its source are aiming at this number specifically.

‍

What a source link has to do before it counts

‍

Not all citations are equal. A usable one meets five conditions.

‍

  • Lands on the exact page, not the document. A 400-page PDF that opens at page 1 is not a citation.
  • Highlights or identifies the passage relied on. Otherwise the reader re-reads the page to find the claim.
  • Names the document type and date. Determines whether the page can prove the point at all.
  • Works for every factual statement, not just some. Partial sourcing hides which statements are unsourced.
  • Survives export into the claim file or a report. An answer that cannot be shown later cannot be relied on later.

Structured claim file review for bodily injury is judged on these mechanics rather than on the fluency of the summary. A well-written answer with weak sourcing is worse than a plain one with strong sourcing.

‍

Regulatory documentation expectations

‍

Verification is not only an internal quality question. It is a documented expectation.

‍

The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers was adopted on 4 December 2023. It sets out state regulators' expectations for insurers using AI systems. It asks for a written AI systems programme, governance, testing and validation, and oversight of third parties acting for the insurer.

‍

It also tells insurers what may be requested later. The bulletin "advises insurers of documentation that a state Department of Insurance may request during an investigation or examination."

‍

Federal guidance points the same way. NIST AI 600-1, the Generative AI Profile published on 26 July 2024, is a companion to the AI Risk Management Framework. It is voluntary, and it is written to help organisations build trustworthiness into how AI systems are evaluated and used.

‍

Neither document asks whether the vendor promised accuracy. Both ask what the insurer did to check.

‍

How much checking is enough

‍

Verifying every sentence of every answer defeats the purpose. Verifying nothing defeats the file.

‍

A workable split has three tiers.

‍

  • Always check any answer that moves a reserve, supports a coverage position, or goes into a demand response or a large loss notice.
  • Sample routine answers at a fixed rate, and record the rate. A sampling policy you can describe is worth more than an informal habit.
  • Never rely on an unsourced answer for anything external, however plausible it reads.

Log what was checked. The log is what turns individual diligence into a programme, and it is what an examiner or opposing counsel will ask to see.

‍

What this does not fix

‍

Source checking catches wrong and misgrounded answers. It does not catch three other things.

‍

It does not catch a complete file that is missing a provider nobody requested. It does not catch a question nobody thought to ask. And it does not catch a judgement call that was documented correctly and made badly.

‍

Those stay with the adjuster, which is where they belong. Verification tells you the facts in front of you are real. Deciding what they mean is still the job.

‍

Key takeaways

‍

  • An AI answer fails in two ways. It states something false, or it cites a real page that does not support the claim. The second form is harder to spot, because the page number and the document are both real.
  • Retrieval does not remove the problem. A Stanford Law School study found leading retrieval-based legal research tools hallucinate between 17% and 33% of the time.
  • Run four questions against any answer that changes a decision. Does the cited page exist in this file? Does it say what the answer says? Is it the right document type to prove that? Does anything else in the file contradict it?
  • A citation only counts if it lands on the exact page, identifies the passage, names the document type and date, covers every factual statement, and survives export.
  • In a demo, test questions you can already answer, ask for something that should not exist, and time the verification rather than the generation.
  • The NAIC Model Bulletin, adopted 4 December 2023, tells insurers that a Department of Insurance may request AI governance and testing documentation during an examination.

‍