Ask anyone shopping for AI software this year for a statistic and you will get the same one back without fail. Ninety five percent of AI pilots fail. It comes from MIT. Say it in a vendor call or a board meeting and watch every head in the room nod at once.
Almost nobody nodding has opened the report the number comes from.
A working paper became a fact
The number traces to one document, The GenAI Divide, State of AI in Business 2025, published by MIT's Project NANDA in July 2025. It is not a peer reviewed study. Its own cover page calls the results preliminary findings, reviewed by exactly one person, built on structured interviews with 52 organizations, survey responses from 153 senior leaders, and a review of more than 300 public AI initiatives, most of them gathered at four industry conferences.
That is a real piece of research, and a modest one. It became a fact of the industry three weeks later, when Fortune covered it under a headline built around the one number that would travel: 95 percent. From there it moved into decks, sales calls, and posts written by people who had read the headline and nothing underneath it.
The number does not say what gets repeated
Here is what the report actually measured. Ninety five percent of the organizations in its sample were getting zero measurable return on generative AI. That is a claim about business outcomes, not a claim that 95 percent of pilots fail outright.
The narrower, more alarming version, the one about pilots specifically, applies to one slice of the data: custom, enterprise grade tools built or bought for a particular workflow. Of that group, 60 percent of organizations evaluated such a tool, 20 percent reached a pilot, and 5 percent reached production. Run the arithmetic on the group that actually ran a pilot, and about one in four made it to production, not one in twenty. That is a rough quarter, a bad number for anyone selling custom AI tooling. It is not a 95 percent failure rate.
The report's other finding almost never survives the retelling. Generic tools like ChatGPT and Copilot told a different story: over 80 percent of organizations explored or piloted them, and nearly 40 percent report actual deployment. A healthy adoption curve was sitting in the same document as the catastrophe, and the catastrophe was the version that traveled.
The citation kept drifting after the fact
The sample size did not survive the retelling either. The report states 52 organizations interviewed and 153 leaders surveyed. Fortune's own coverage described 150 interviews and a survey of 350 employees, figures that do not appear anywhere in the document. Nobody corrected this loudly enough for the correction to travel as far as the error did.
By this year the drift has compounded again. Search for the statistic today and you will find a small industry of posts confidently attributing different numbers, 86 percent here, 88 somewhere else, 89 in a third place, each one credited to a named research firm with total certainty and, in most cases, no traceable source document behind it. The number has stopped being reported. It is being regenerated, each retelling producing a plausible variant of the last one with nobody checking it against anything.
One more detail rarely makes it into the retelling. Project NANDA is not a neutral survey house. It is a research group building infrastructure and protocols for distributed AI agents, the exact category its own report warns enterprises about approaching carelessly. That does not make the research dishonest. The authors flagged their own findings as preliminary and anonymized their sources, which is more caution than most of what got built on top of it. But it belongs in the room whenever the number gets cited, and it has dropped out of every retelling this piece could find.
Verification is the scarce resource, not the statistic
None of this means enterprise AI is going well. It might be exactly as troubled as the sound bite claims, on different evidence than the sound bite is actually standing on. That is the problem. An entire industry's working number for how badly AI projects fail rests on a document almost nobody in that industry has read, describing a narrower population than the one the number gets applied to, with a sample size that already changed once between the report and the headline and has kept drifting since.
The report itself was honest about its limits. Its own methodology page called the findings preliminary. It was everyone downstream, the headline writers, the deck builders, the people repeating it in meetings, who dropped every qualifier on the way to a round number that was scary enough to remember.
A statistic that has traveled this far, picking up a new sample size and a new attributed source at every stop, is not evidence anyone can build a decision on. Before repeating the next frightening number about AI, or the next reassuring one, find the document it came from and read what it actually says. Almost nobody bothers, which is exactly why doing it is worth something.