Higher Agency · SF/Berlin/Remote
·
Insight · Operator note

Nobody checked the 95%

The most-quoted statistic in enterprise AI comes from 153 survey responses, describes something other than failure, and is now used to end conversations it should start.

By the founder · Higher Agency

Adding AI features hasn't moved the number on the spreadsheet. Then someone in the room says ninety-five percent of enterprise AI pilots fail, and the meeting relaxes. The statistic explains the disappointment. It also ends the conversation, and ending the conversation is the expensive part.

We are not going to tell you the number is fake. It isn't. We are going to show you where it came from, because we think you will draw a different conclusion once you have seen it, and because a number nobody has checked is a poor thing to set a budget against.

It comes from a report called The GenAI Divide: State of AI in Business 2025, published in July 2025 by Project NANDA, a research initiative at the MIT Media Lab. Twenty-six pages. The sentence underneath the headline reads: "Just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact."

Three things in that sentence are not what the headline says.

The 95% was never measured. It is what remains after the 5%, which means nobody counted 95 failures. They counted a small group of clear winners and called everything else the rest. Second, "no measurable P&L impact" is not the same claim as "failed." A project can work, be liked, save real hours, and still not show up in a quarterly line item, particularly if nobody set out to make it show up. Third, the sentence is about integrated pilots, a subset, not about every AI project a company runs.

Then there is the evidence base: 52 executive interviews, 153 survey responses, and a review of more than 300 publicly disclosed AI initiatives. The survey respondents were recruited at four conferences, which is a group of people who travel to hear about AI rather than a cross-section of the economy. The report describes its own findings as preliminary and lists what it does not cover, including regional breakdowns, vendor benchmarks, and case studies.

One more thing worth knowing. Project NANDA builds agentic AI infrastructure, and the report closes by making the case for it. That does not make the work wrong. It does mean the finding and the product point the same direction, which is a fact you would want disclosed if a vendor handed you the same slide.

The criticism has been specific. Futuriom went looking for the passage where the 95% is actually derived, reported that it could not find one, and asked for the supporting data or a retraction. We can find no evidence that MIT withdrew or corrected the report, and we are not going to imply otherwise. The number stands. It is simply softer than the way it gets quoted.

Here is the part that got dropped on the way to the headline. The same report found employees at more than nine in ten companies already using AI at work, mostly personal tools they picked up without waiting for a program, against roughly 40% of companies that had bought an official subscription. Fortune ran that finding too, one day after the story you have heard, under a headline about a booming shadow AI economy. Same report, same outlet, consecutive days, opposite weather. VentureBeat went further and argued the report had been misread outright.

Read whole, the report says something more useful than "AI does not work." It says AI is already running inside almost every company, and almost nobody can prove what it paid for.

That is a measurement problem sitting in front of a technology problem. It is also the more actionable of the two, because measurement is decided before the build, by people in a room, in an afternoon.

The uncomfortable version, from the pilots we are brought in to diagnose: a good number of them could not have demonstrated success even if they had worked perfectly. No baseline was captured. No owner was named. Nobody agreed in advance what would count. When the review came, the team had anecdotes and the finance seat had a spreadsheet, and the spreadsheet won, as it always does.

So we do the boring thing first. Before anything is built, we agree the number, who owns it, what the baseline is today, and the date we look. Written down, before scoping ends. It is the least impressive part of the engagement and the part that decides whether the work survives its first review.

This is the same standard we hold our own systems to. Every number a system of ours produces carries a tag saying where it came from: sourced, estimated, modeled, or derived. A surface that cannot say where its numbers come from is not finished. We apply that to a statistic in a board deck for the same reason we apply it to an agent's output, which is that an unsourced number is a confident guess wearing better clothes.

For contrast, here is a number we do use. S&P Global found that 42% of enterprises abandoned the majority of their AI initiatives in 2025, up from 17% the year before. We quote it because you can see who counted, what they counted, and what changed year over year. It is a worse headline and a better fact.

If your pilot flattened, the honest question is not whether you are inside somebody's 95%. It is whether anyone agreed, before the work started, what moving the number would have looked like. That is a short conversation, and it is the one we would start with.

Next step

What would have counted as working?

Thirty minutes with a founder. Tell us what the pilot was supposed to move, and we'll tell you whether it was ever measurable.

Talk to a founder