July 21, 2026
3 minute read. 643 words at 250 a minute, counting the opening, the section paragraphs, the call and the Monday morning list. Component notes, sources, disclosures and corrections are not counted. Scores on this page as of August 9, 2026. Any change is logged below.
VerdictHalf True
Hype Index53 / 100
SourceMIT Project NANDA
Call resolvesJuly 2027
The record No call has reached its resolution date yet. The first is due January 2027. Every call is logged with a date, and the misses get published the same way the hits do.

A CIO forwarded me a board deck last month. Slide 4 had one number on it, in 90 point type, with no footnote. His board had already voted on it.

The claim

“95% of enterprise AI pilots fail.”

Quoted in board decks, keynote slides, and vendor pitches since August 2025. Usually with no source attached.

The research is real. The number is real. What people do with it is not.

Real finding, wrong scope. Correct in part, misapplied in practice.

The verdict
Holds upAll hype
53
HALF TRUE
HYPE INDEX · 0 TO 100
52
Organizations interviewed. That is the sample carrying a statistic now quoted at billion-dollar budget meetings.

Where the 95% actually comes from

The number never measured what it is used to prove.

The source is The GenAI Divide: State of AI in Business 2025, published July 2025 by MIT's Project NANDA.

Page two of the report labels it Preliminary Findings. The methodology is 52 structured interviews, 153 survey responses collected at four industry conferences, and a review of 300 publicly disclosed AI initiatives.

The 95% measures the inverse of a single funnel in section 3.2, for custom and task specific GenAI tools only. Sixty percent investigated. Twenty percent piloted. Five percent reached production.

General purpose tools were not in that funnel. The report says elsewhere those convert at roughly 83%.

Directly beneath the exhibit, the authors write that the figures are “directionally accurate based on individual interviews rather than official company reporting.”

They also define success as a deployment that users or leaders “remarked as” causing sustained productivity or P&L impact. Remarked. Not audited. Not measured.

The observation window was six months. In the appendix, the authors note this “may be insufficient” and could be “understating success rates.”

None of that survives the trip to a slide.

Three things worth knowing

The report is preliminary, self-reviewed, and published by a party with a stake in the answer.

The reviewer is one of the authors. Page two lists a single reviewer, Pradyumna Chari, Project NANDA. Page one lists Pradyumna Chari as a co author. This is not peer review. It was never presented as peer review.

The report contradicts itself on its own second most quoted figure. Section 3.4 puts sales and marketing at “approximately 70 percent” of GenAI budget. Section 6.3 puts it at 50%. Both appear in the same document.

The publisher has a position. The appendix states that Project NANDA “builds on Anthropic's Model Context Protocol and the Google/Linux Foundation A2A to create infrastructure for distributed agent intelligence at scale.” The report's conclusion is that the fix is agentic systems with persistent memory. That happens to be the category NANDA builds.

That does not make the work dishonest. It makes it interested. Interested research can still be correct. It just should not be quoted as a neutral scoreboard.

What holds up

Strip the headline and the findings underneath are worth more than the number that made it famous.

Ninety percent of surveyed employees use personal AI tools for work. Forty percent of their companies bought a subscription. That gap is the real finding and almost nobody quotes it.

Externally sourced tools reached deployment about 67% of the time. Internal builds, about 33%. Twice the success rate for buying over building.

Mid market firms went from pilot to production in about 90 days. Enterprises took nine months or more.

Half to seventy percent of budget went to sales and marketing, while the documented savings sat in the back office. Two to ten million a year from eliminating outsourced customer service and document processing.

Read that way, the report is not a verdict on AI. It is a verdict on how organizations buy it.

What would change our mind

We are telling you in advance what evidence would move this score.

Release of the underlying data. A replication with a defined, audited success metric and a window longer than six months. Either would move this score.

How this scored

Five components, twenty points each. Higher means more hype risk.

Source quality7 /20
Rubric 5 to 9: Preprint, working paper, or an official statistical release that has not been reviewed. The report is self-published and not identified as peer reviewed, although it discloses a multi-method methodology covering 300 initiatives, 52 organizations, and 153 senior leaders.
Sample and method challenged12 /20
Rubric 10 to 14: Adequate size, self-selected or convenience sample, method disclosed. The sample is substantial and the method is described, but survey respondents came from four conferences, company data was anonymized, and the underlying data is not provided for inspection.
Independence6 /20
Rubric 5 to 9: Publisher has an indirect interest, offset by independent co-authors or external review. The evidence does not show that NANDA sells the recommended remedy, though its program is financially supported by AGNI research projects and member organizations, creating some indirect interest.
Replication13 /20
Rubric 10 to 14: Not reproduced. Supporting evidence exists but comes from interested parties or a different method. No independent reproduction is provided; the evidence contains only the original report and an archive snapshot of it.
Drift15 /20
Rubric 15 to 17: Applied well outside the measured population, conditions, or timeframe. The source says 95% of organizations got zero return and that just 5% of integrated pilots extracted value, while circulation converts this into the broader and stronger claim that 95% of enterprise AI pilots fail.

challenged marks a component that a reader formally disputed and where the dispute was upheld, moving the score. Read the challenge and its outcome.

The full method and the rubric. If you think a component sits in the wrong band, say which band and why: challenges are published with their outcome.

Who scored this

The score above is the council's. Five models from five laboratories read the fetched text of the sources listed below, score the claim against the same published rubric independently, then review each other anonymised. Each edition is scored three times, because one run is a sample of size one, and the five component scores are the median of those runs.

Council score53 / 100, Half True
Across 3 runs57, 51, 53
Members apart34 points
Editor scored it69 / 100, Overstated

Where a person landed. Before the council took over scoring, the editor read the same sources and put this at 69, Overstated. The council came in 16 points lower, in a different band. Both readings are kept and dated, and both get graded against this edition's call when it resolves. That comparison is the point of keeping the old number rather than deleting it.

The three runs moved 6 points. Close enough to publish the median, far enough that the range belongs next to it. The verdict held in every run. The council is steady and the members are not: 34 points separate the highest and lowest scorer below. A stable median built out of wide disagreement is a different object from a stable median everyone agrees on, and the individual scores are printed so you can tell them apart.

Model Score Verdict In one line
GPT-5.6 57 Half True The claim reflects a real but imperfectly measured finding about zero-return organizations or integrated pilots, overstated as a universal failure rate for enterprise AI pilots.
Claude Opus 4.8 59 Half True A single preliminary MIT NANDA report saying 95% of organizations saw zero GenAI return has been reworded into '95% of pilots fail' and circulated sourceless in vendor decks.
Grok 4.5 63 Overstated A real but narrowly scoped NANDA finding has been stripped of its organizational and 'zero-return' framing and turned into a free-floating failure statistic.
Sonar Pro 29 Holds Up The claim is broadly supported by one MIT report but is being repeated more broadly than the evidence cleanly justifies, especially when treated as a universal failure rate for all enterprise AI pilots.
DeepSeek 45 Half True The claim is based on a reputable but non-peer-reviewed study with a reasonable methodology, but it is often cited without its original context and qualifiers.
Opinion of the council · written by GPT-5.6

The claim reflects a real but imperfectly measured finding about zero-return organizations or integrated pilots, overstated as a universal failure rate for enterprise AI pilots.

  • Source quality. The report is self-published and not identified as peer reviewed, although it discloses a multi-method methodology covering 300 initiatives, 52 organizations, and 153 senior leaders.
  • Sample and method. The sample is substantial and the method is described, but survey respondents came from four conferences, company data was anonymized, and the underlying data is not provided for inspection.
  • Independence. The evidence does not show that NANDA sells the recommended remedy, though its program is financially supported by AGNI research projects and member organizations, creating some indirect interest.
  • Replication. No independent reproduction is provided; the evidence contains only the original report and an archive snapshot of it.
  • Drift. The source says 95% of organizations got zero return and that just 5% of integrated pilots extracted value, while circulation converts this into the broader and stronger claim that 95% of enterprise AI pilots fail.

The author is not chosen. It is whichever member landed closest to the consensus.

Dissent · Sonar Pro, scored 29

A stronger case for a lower score is that the report itself explicitly states the 95% figure and is based on a named, multi-method study, so the popular shorthand may be only a mild simplification rather than hype.

Printed verbatim. The dissenter is whichever member landed furthest from the consensus, so the strongest objection reaches you whether or not it is convenient.

Publisher's note. Nothing here is edited. The council's number is whatever the council produced, including where it disagrees with the published score and where the members disagree with each other. Agreement between models is not the same thing as accuracy, which is why the spread sits next to the median rather than under it. The scoring code is open source at github.com/marklynd/quorum, and every run writes a dated transcript before the call resolves. How the council runs, and what it is for.

The call · logged July 21, 2026 · resolves July 2027

No peer reviewed replication will put enterprise AI ROI failure at 95%. When better studies land, the honest range will fall between 60 and 80 percent. Most of that spread will come from how each study defines the word success.

How it resolves. Resolves HELD if, by July 21, 2027, no peer reviewed study of enterprise AI ROI reports a failure rate at or above 90%. Resolves MISSED if one or more does. Resolves VOID only if no qualifying study is published, in which case it carries forward one year and is recorded as unresolved rather than quietly dropped.

This is a judgment call, and it goes on the record either way.

Monday morning

  1. Find out who is quoting the 95% inside your organization and ask them for the source. Most cannot produce it. That conversation is more valuable than the statistic.
  2. Count your own funnel. How many AI pilots have you started, how many are in production, and who decided what production means. If nobody owns that definition, your internal number is worth exactly as much as the one on the slide.
  3. Look at what your people already use without asking. The 90% versus 40% gap is the cheapest signal you will get about which workflows actually want AI.

Sources

Second reader. 21 of 21 printed figures traced to 3 linked sources, checked July 26, 2026. Nothing printed here is missing from the documents below. How this runs.

Every figure above traces to one of these. Where a number carries a date, use the date.

Corrections

July 26, 2026. Re-scored against rubric v1.0 when it was published. Source quality 14 to 13, Sample and method 15 to 14, Independence 16 to 12, Replication 15 to 12, Drift 14 to 16. Total moved from 74 to 67. Publication date corrected from July 26 to July 21, the Tuesday it went out. The source line now states that the report PDF is a third-party mirror rather than a publisher-hosted copy.

July 26, 2026. Challenge C-004 upheld. Sample and method moved from 14 to 16. Fifty-two interviews is a small sample for a claim about enterprise AI generally, which is the 15 to 17 band, and scoring it as adequate size contradicted this edition's own headline statistic. Total moved from 67 to 69. The verdict is unchanged at Overstated.

July 27, 2026. Rescored. This edition was originally scored 69 out of 100, Overstated, by the editor. Scoring moved to a council of five models on 2026-07-27, and the council put the same claim, read from the same sources, at 53, Half True. The council score is now the published one. Total moved -16 points. The original score is kept on the record and both readings will be graded against this edition’s call when it resolves.

Every change is logged here rather than edited away. The policy.

Cite this

The Hype Index, July 21, 2026. “95% of enterprise AI pilots fail.” scored 53/100, Half True. Primary source: The GenAI Divide: State of AI in Business 2025, MIT Project NANDA, July 2025. Scored by a council of five models, median of 3 runs, range 6; published by Mark Lynd. https://thehypeindex.com/edition/95-percent-of-enterprise-ai-pilots-fail/

Everything in that line is checkable. That is the point of it. Think something here is wrong? Challenge the score.

Score it yourself The checker opens loaded with this edition's five component scores. Move the one you disagree with and see how much of the verdict it was carrying. It runs in your browser and nothing is sent anywhere.
Forward this
"95% of enterprise AI pilots fail." — scored 53/100, Half True, by The Hype Index.
LinkedIn X Email

Every edition is public. No paywall, no registration to read.

The record

ClaimVerdictIndexCall resolves
Security keys stop 100% of phishing.
August 6, 2026
Half True 43 January 2027
There are 4.8 million unfilled cybersecurity jobs worldwide.
August 4, 2026
Overstated 76 January 2027
There is a plastic spoon's worth of microplastic in your brain, and it is causing dementia.
July 30, 2026
Half True 53 January 2027
Superhero movies are dead.
July 28, 2026
Half True 50 January 2027
95% of enterprise AI pilots fail.
July 21, 2026
Half True 53 July 2027
Get every verdict

Two editions a week. One scored claim each.

Tuesday and Thursday at 1:00 PM Central. Every call is dated and graded in public, including the ones we get wrong.

Thousands of readersFree foreverUnsubscribe anytime

One email, one scored claim, no pitch. We never sell or share your address.