Edition 001 · July 26, 2026
The record This is edition 001. No calls have resolved yet. Every one we make is logged with a resolution date, and the misses get published the same way the hits do.

A CIO forwarded me a board deck last month. Slide 4 had one number on it, in 90 point type, with no footnote. His board had already voted on it.

The claim
“95% of enterprise AI pilots fail.”

Quoted in board decks, keynote slides, and vendor pitches since August 2025. Usually with no source attached.

The research is real. The number is real. What people do with it is not.

Sound evidence, carried far past what it actually measured.

The verdict
Holds upAll hype
74
OVERSTATED
HYPE INDEX · 0 TO 100
52
Organizations interviewed. That is the sample carrying a statistic now quoted at billion-dollar budget meetings.

Where the 95% actually comes from

The number never measured what it is used to prove.

The source is The GenAI Divide: State of AI in Business 2025, published July 2025 by MIT's Project NANDA.

Page two of the report labels it Preliminary Findings. The methodology is 52 structured interviews, 153 survey responses collected at four industry conferences, and a review of 300 publicly disclosed AI initiatives.

The 95% is not an ROI measurement across enterprise AI. It is the inverse of one funnel in section 3.2, for custom and task specific GenAI tools only. Sixty percent investigated. Twenty percent piloted. Five percent reached production.

General purpose tools were not in that funnel. The report says elsewhere those convert at roughly 83%.

Directly beneath the exhibit, the authors write that the figures are “directionally accurate based on individual interviews rather than official company reporting.”

They also define success as a deployment that users or leaders “remarked as” causing sustained productivity or P&L impact. Remarked. Not audited. Not measured.

The observation window was six months. In the appendix, the authors note this “may be insufficient” and could be “understating success rates.”

None of that survives the trip to a slide.

Three things worth knowing

The report is preliminary, self-reviewed, and published by a party with a stake in the answer.

The reviewer is one of the authors. Page two lists a single reviewer, Pradyumna Chari, Project NANDA. Page one lists Pradyumna Chari as a co author. This is not peer review. It was never presented as peer review.

The report contradicts itself on its own second most quoted figure. Section 3.4 puts sales and marketing at “approximately 70 percent” of GenAI budget. Section 6.3 puts it at 50%. Both appear in the same document.

The publisher has a position. The appendix states that Project NANDA “builds on Anthropic's Model Context Protocol and the Google/Linux Foundation A2A to create infrastructure for distributed agent intelligence at scale.” The report's conclusion is that the fix is agentic systems with persistent memory. That happens to be the category NANDA builds.

That does not make the work dishonest. It makes it interested. Interested research can still be correct. It just should not be quoted as a neutral scoreboard.

What holds up

Strip the headline and the findings underneath are worth more than the number that made it famous.

Ninety percent of surveyed employees use personal AI tools for work. Forty percent of their companies bought a subscription. That gap is the real finding and almost nobody quotes it.

Externally sourced tools reached deployment about 67% of the time. Internal builds, about 33%. Twice the success rate for buying over building.

Mid market firms went from pilot to production in about 90 days. Enterprises took nine months or more.

Half to seventy percent of budget went to sales and marketing, while the documented savings sat in the back office. Two to ten million a year from eliminating outsourced customer service and document processing.

Read that way, the report is not a verdict on AI. It is a verdict on how organizations buy it.

What would change our mind

We are telling you in advance what evidence would move this score.

Release of the underlying data. A replication with a defined, audited success metric and a window longer than six months. Either would move this score.

How this scored

Five components, twenty points each. Higher means more hype risk.

Source quality14 /20
Self-published report labeled Preliminary Findings. Not peer reviewed and never presented as such.
Sample and method15 /20
52 interviews, 153 conference surveys, 300 public initiatives. Underlying data not released.
Independence16 /20
The publisher builds infrastructure in the exact category the report recommends as the fix.
Replication15 /20
No independent replication with a comparable method.
Drift14 /20
The headline is used as an ROI verdict on all enterprise AI. The source measured one narrow funnel.

The full method. If you think a component is wrong, the argument is specific, and that is the point.

The call · logged July 26, 2026 · resolves July 2027

No peer reviewed replication will put enterprise AI ROI failure at 95%. When better studies land, the honest range will fall between 60 and 80 percent. Most of that spread will come from how each study defines the word success.

How it resolves. Resolves HIT if, by July 26 2027, no peer reviewed study of enterprise AI ROI reports a failure rate at or above 90%. Resolves MISS if one or more does. Resolves VOID only if no qualifying study is published, in which case it carries forward one year and is recorded as unresolved rather than quietly dropped.

A judgment, not a finding. It goes in the record either way.

Monday morning

  1. Find out who is quoting the 95% inside your organization and ask them for the source. Most cannot produce it. That conversation is more valuable than the statistic.
  2. Count your own funnel. How many AI pilots have you started, how many are in production, and who decided what production means. If nobody owns that definition, your internal number is worth exactly as much as the one on the slide.
  3. Look at what your people already use without asking. The 90% versus 40% gap is the cheapest signal you will get about which workflows actually want AI.

Source

The GenAI Divide: State of AI in Business 2025. MIT Project NANDA, July 2025. Every figure above is drawn from the report text, including its stated methodology, research limitations, and appendix.

The record

No.ClaimVerdictIndexCall resolves
001 95% of enterprise AI pilots fail.
July 26, 2026
Overstated 74 July 2027
Get every verdict

Two editions a week. One scored claim each.

Tuesday and Thursday at 1:00 PM Central. Every call is dated and graded in public, including the ones we get wrong.

Free foreverTuesdays & ThursdaysUnsubscribe anytime