The 95% number is doing a lot of work
Everyone quotes the statistic that 95% of AI pilots fail. Four widely cited figures range from 40% to 95% — because they measure four different things. A look at what the numbers actually say, from someone who benefits from them.
You have seen the number. Ninety-five percent of generative AI pilots deliver no measurable return. It has been in every newsletter, every conference keynote, and — I should admit this early — it is a very convenient number for someone who sells AI delivery consulting. It says the problem is enormous, and it says you need help.
That is exactly why it deserves a harder look than it usually gets.
Where it comes from
The figure originates with MIT Media Lab's Project NANDA. The underlying work is roughly 150 interviews with leaders, a survey of about 350 employees, and an analysis of 300 publicly disclosed AI deployments. For a qualitative study that is a reasonable base. For a statistic quoted as though it were a census of the global economy, it is thin.
The precise claim is also narrower than the way it travels. NANDA reports that 95% of pilots show no measurable impact on profit and loss. That is not the same sentence as “95% of AI projects fail,” but that is the sentence people repeat.
Four numbers that do not agree
Put the commonly quoted figures side by side and the problem becomes visible immediately.
Percentage quoted, and what each one actually counts
agentic AI projects it expects to be cancelled by end of 2027
AI projects that fail — about twice the rate of conventional IT
AI pilots that never reach production
GenAI pilots with no measurable impact on profit and loss
| Source | Figure | What it measures | Measured or forecast |
|---|---|---|---|
| Gartner | 40% | agentic AI projects it expects to be cancelled by end of 2027 | forecast |
| RAND | 80% | AI projects that fail — about twice the rate of conventional IT | Measured or forecast |
| IDC | 88% | AI pilots that never reach production | Measured or forecast |
| MIT Project NANDA | 95% | GenAI pilots with no measurable impact on profit and loss | Measured or forecast |
A range from 40% to 95% is not measurement error. It is four organisations answering four different questions. Gartner is counting projects it expects to be cancelled — and it is forecasting, based on a mid-2025 poll of more than 3,400 organisations, not reporting. RAND counts projects that failed. IDC counts pilots that never reached production. NANDA counts pilots with no P&L impact.
Those are four genuinely different outcomes. A pilot can reach production and still not move the P&L. A pilot can be cancelled and have been the correct decision. Collapsing all of them into “AI fails” throws away the only information that would help you.
Failure is not the useful category
Here is the part that bothers me most about the way the number is used. A pilot that spends six weeks and €40,000 establishing that a use case does not work is not a failure. It is cheap information, arrived at in the correct order. It shows up in every one of these statistics as a loss.
The expensive failure is not the pilot that stops. It is the pilot that keeps going because nobody defined what would make it stop.
I have seen more money burned on projects that could not be killed than on projects that were killed early. None of the headline statistics distinguish between the two, which makes them close to useless as a guide to your own decisions.
What the studies do agree on
The numbers diverge. The diagnoses converge — and that convergence is the part worth your attention. Across RAND, IDC, Gartner and NANDA, the recurring causes are the same, and none of them are about model quality:
- No measurable business objective attached to the work on day one.
- Data readiness, which is the largest single driver and is almost always discovered late — after the pilot has already promised something.
- Weak integration into the actual workflow, so the output lands somewhere nobody is obliged to look.
- Executive sponsorship that fades once the demo stops being novel.
- No named owner for the output after handover, as distinct from an owner for the project.
Four studies that cannot agree on a percentage agree completely on the causes. That tells you the causes are real and the percentage is decoration.
The counter-evidence nobody quotes
It is worth noting that the NANDA report has been criticised on exactly these grounds. Futuriom, which pushed back hardest, maintains a database of more than 130 documented enterprise AI deployments with named organisations and named results. Deloitte's own survey found around a quarter of organisations had already moved 40% or more of their pilots into production. Those findings get a fraction of the circulation, because “a meaningful minority is doing fine” is not a headline.
What to do with any of this
Stop asking what the failure rate is. It is not a number that applies to you. Ask the four questions the studies actually converge on, before the pilot starts:
- What number on which report changes, and by how much, if this works?
- Who owns that number today — before the project exists, and after it ends?
- What would have to be true for us to stop, and who is allowed to call it?
- Has anyone looked at the actual data yet, or are we assuming it is usable?
If you cannot answer all four, your pilot will join whichever statistic you find most persuasive. If you can, the industry failure rate stops being relevant to you, which was always the point.
And apply the same suspicion to me. I sell the fix for the problem this number describes. That is a reason to check my reasoning, not to accept my framing.
Sources
- AI delivery
- Governance
- Measurement