
MIT found 95% of enterprise generative AI pilots fail to deliver measurable impact. Here's why claims AI pilots stall, and how a Proof of Value framework produces a decision instead.

About 95% of enterprise generative AI pilots fail to deliver measurable business impact, according to MIT's 2025 “GenAI Divide” study, and separate research from IDC found that for every 33 AI proofs of concept a company launches, only four reach production, an 88% failure rate at the scaling stage [1][2][4]. The technology is usually not the problem. Pilots die during integration into real workflows, on data that is messier in production than in a demo, and on evaluations that never defined what “working” would actually mean. In claims, where AI document analysis is now one of the most heavily piloted use cases in the industry, that last failure point is the most common one, and it is entirely fixable.
The scale of the problem shows up consistently across independent studies published in 2025 and 2026:
These numbers span every industry, not just insurance, which is itself informative: the failure pattern is structural, not sector-specific.
MIT's research concluded that model quality was rarely the blocking issue. Pilots stalled on integration into real workflows, organizational readiness, and data that proved far messier in production than in the demo. A generic chatbot applied to a trivial task can reach 83% adoption quickly, because the task requires little context. The moment a workflow demands real customization and context (which describes almost every claims use case), adoption stalls unless the system can retain feedback and adapt over time [1][3].
Claims organizations describe a familiar pattern. A vendor runs its software on a sample of files, someone reviews the output, the demo looks good, and the team agrees it “works.” Then the carrier tries to turn that impression into a go-forward decision and cannot, because “it works” was never defined precisely enough to support a yes or a no. Does it work well enough to change how adjusters handle files? On which lines? Measured how, against what baseline?
The proof of concept produced a feeling of success without the evidence to act on it, and the project stalls in the gap between the demo and the rollout, often the single most expensive phase of the whole process, because it consumes adjuster time and executive attention and frequently ends with no decision at all.
The two tests answer different questions, and confusing them is what traps many pilots. A proof of concept asks a technical question: does the software run on our files and produce plausible output?
A Proof of Value asks an outcome question: measured against targets we set in advance, is the solution producing the result we are paying for? The technical check can pass while the outcome question goes unanswered, which is exactly where many stalled pilots are.

Carriers who get past the pilot stage share a common discipline: they define what success means, in numbers, before the test begins, and they measure against those numbers when it ends. A working design includes the following steps:
Even a well-intentioned test can fail to settle the decision. The most common mistakes are a baseline that was never measured, so the tool's numbers have nothing to compare against; a curated file set that hides how the tool performs on the real distribution of work; success criteria left loose or renegotiated after seeing the results, so the test stops measuring anything; governance treated as a later step, so a favorable result later fails a security review and the evaluation has to restart; and no plan for what a passing result actually triggers, so a successful pilot loses momentum while the organization works out integration and training. None of these are technology problems. They are design problems, entirely within the organization's control.
What percentage of AI pilots fail?
MIT's 2025 GenAI Divide study found that about 95% of enterprise generative AI pilots fail to deliver measurable business impact. IDC research found a similar pattern at the scaling stage: for every 33 AI proofs of concept a company starts, only about four reach production, an 88% failure rate [1][2][4].
What is the difference between a proof of concept and a proof of value?
A proof of concept tests whether software runs on your files and produces plausible output: a technical question. A Proof of Value tests whether the solution delivers the specific, measurable outcome the organization defined in advance, against a baseline: an outcome question. A tool can pass a proof of concept and still fail to deliver real value, which is why the distinction matters for a purchase decision.
Why do enterprise AI pilots fail to reach production?
MIT's research found the model itself is rarely the blocking issue. Pilots typically fail because of integration difficulty with real workflows, organizational readiness gaps, and production data that is messier and more varied than the clean sample used in the demo. In claims specifically, pilots also fail because teams never define success precisely enough before testing to support a clear go/no-go decision.
How do you measure ROI on an AI pilot?
Define specific, measurable success criteria before the pilot starts, tied to a documented baseline of current performance, rather than after seeing results. In claims document analysis, useful metrics include time to a decision-ready understanding of a file, the share of known findings the tool catches against an answer key, citation fidelity, coverage across file types and lines, and, where the data supports it, outcome measures like reserve accuracy or settlement timing against a control group.
This is the gap amaise built its Proof of Value framework to close: a structured test, run on a carrier's own production files against targets agreed in advance, that ends in a decision rather than another meeting. To design a Proof of Value for your claims organization, contact amaise at hello@amaise.com.
[1] MIT, Project NANDA. “The GenAI Divide: State of AI in Business 2025.” https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
[2] Fortune. “MIT report: 95% of generative AI pilots at companies are failing.” August 18, 2025. https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/
[3] Forbes. “MIT Finds 95% Of GenAI Pilots Fail Because Companies Avoid Friction.” August 26, 2025. https://www.forbes.com/sites/jasonsnyder/2025/08/26/mit-finds-95-of-genai-pilots-fail-because-companies-avoid-friction/
[4] SoftwareSeni. “Why 88 to 95 Percent of Enterprise AI Pilots Never Reach Production.” (citing IDC/Lenovo research) https://www.softwareseni.com/why-88-to-95-percent-of-enterprise-ai-pilots-never-reach-production/
[5] Legal.io. “MIT Report Finds 95% of AI Pilots Fail to Deliver ROI, Exposing the 'GenAI Divide.'” https://www.legal.io/blog/5719519/MIT-Report-Finds-95-of-AI-Pilots-Fail-to-Deliver-ROI-Exposing-GenAI-Divide