ELL ADVISORY

The 95% Failure Rate Is Mostly a Measurement Failure

Fawad Bhatti, Founder of Ell Advisory
Founder, Ell Advisory · Ex-Hilti Principal PM · HEC Paris MBA
16 min read

TL;DR

The "95% of AI pilots fail" statistic comes from MIT Project NANDA's The GenAI Divide: State of AI in Business 2025, published July 2025. Its success criterion is the part nobody quotes: "deployment beyond pilot phase with measurable KPIs, with ROI impact measured 6 months post-pilot." Read that carefully. A project that delivered real benefit but never captured a before-state cannot pass — it is counted as a failure by construction. In the UK mid-market that describes most projects I see. The technology worked; nobody wrote down what things looked like beforehand. Six steps to fix that below, and the Admin Tax Index gives you the hours-to-money conversion. Start with where your admin time actually goes.

A managing director told me in June that he was not going to do any more AI projects, because "95% of them fail anyway." He had read it somewhere. He was quite certain about it, and he was using it the way people use these numbers — as permission to stop thinking.

I asked him what the 95% were failing at. He did not know. Neither, it turned out, did anyone else in the conversation.

So I went and read the definition. It is one sentence long, and it changes what the number means.

52

Organisations interviewed

Plus 153 leaders surveyed at four conferences

6 months

Window for ROI to appear

Measured post-pilot, adjusted for dept size

"measurable KPIs"

Required for a pilot to count as a success

No baseline, no pass

Jan–Jun 2025

Fieldwork period

Published July 2025, viral that August

What the study actually is

The GenAI Divide: State of AI in Business 2025 was produced by MIT Project NANDA — Networked Agents And Decentralized Architecture — and published in July 2025. It went everywhere in August after Fortune and The Register picked it up.

Its method: a systematic review of 300+ publicly disclosed AI initiatives, structured interviews with 52 organisations, and survey responses from 153 senior leaders gathered at four industry conferences, with fieldwork running January to June 2025.

The report is candid about its own limits. It states that the sample may not represent all enterprise segments or geographies, and that selection bias is possible among organisations willing to take part. It is a version 0.1 document, and its interview data is described as directionally accurate from individual interviews rather than official company reporting.

A note on my own sourcing

The report PDF returned an HTTP 403 to my automated fetch on 29 July 2026, so the methodology above comes from secondary summaries that agree with one another, not from my reading the primary document. I would rather tell you that than imply a level of verification I did not reach. Note also that Fortune describes the sample as "150 interviews with leaders, a survey of 350 employees" — which disagrees with every other source I checked. I could not resolve it, so I am using the corroborated 52 and 153 figures and flagging the conflict. The most-cited AI statistic of 2025 is reported with different sample sizes by different outlets, and almost nobody repeating "95%" has looked at either.

The sentence inside the number

Here is the success criterion, and it is worth reading twice:

Deployment beyond pilot phase with measurable KPIs, with ROI impact measured 6 months post-pilot, adjusted for department size.

Three conditions. Production, not pilot. Measurable KPIs. ROI actually measured, six months on.

Now think about a real mid-market project. A manufacturer automates job sheet entry. Two coordinators who spent most of a day each week re-keying paperwork stop doing it. The work genuinely goes away. Everyone agrees it is better.

Did they define the KPI before starting? Usually not. Did they measure ROI six months later against a recorded before-state? Almost never — by then the team has moved on and nobody remembers precisely how long it used to take.

That project fails conditions two and three. It goes in the 95%. Not because the AI did not work — because nobody wrote down what things looked like beforehand.

Exhibit 1 — The criterion, applied

A real automation that works, scored against the study's definition of success

ConditionTypical UK mid-market projectResult
Deployed beyond pilotYes — the coordinators genuinely stopped re-keying job sheetsPass
Measurable KPIs definedNo metric agreed before go-live. Everyone agrees “it’s much better”Fail
ROI measured 6 months post-pilotNobody recorded the before-state, so there is nothing to measure againstFail

Success criterion as stated in MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025", July 2025. The right-hand column is the author's illustration of a representative project, not data from the report.

I want to be careful here, because it would be easy to overclaim. I am not saying the 95% is fake. The report found genuine problems, and some of them are worth more attention than the headline.

The two most useful, per Fortune's account: buying a specialised vendor solution outperformed building internally by a wide margin — 67% success against 33% — and back-office automation outperformed sales and marketing deployments. That is an uncomfortable finding for anyone selling custom builds, including me, and it is a good argument for being specific about when a build is genuinely warranted.

There are fair criticisms of the study too. Six months is a short window for AI ROI, particularly where the benefit is released capacity rather than removed cost. A pure P&L test ignores efficiency, churn and conversion. The 300 public initiatives were reviewed without a published account of how. And conference-recruited respondents are not a random sample.

But none of that rescues the project in my example. It still failed on measurement, and measurement is the one thing entirely within your control.

Capture the baseline before you buy

Everything below is my own method rather than a published standard, so weigh it accordingly. It costs about two weeks of light effort and it is the difference between a project you can defend and one you cannot.

Six steps, in order

Capture the before-state before you buy

01

Not after the pilot, not during procurement — before. Once the tool is in, the before-state is unrecoverable, and memory is not a baseline. Two weeks of someone counting is enough. This is the step people skip and it is the only one that cannot be done later.

Count hours first, convert to money second

02

Hours are observable and difficult to argue with. Money is derived, and the derivation is where every argument happens. Keep them separate so a finance director can challenge the conversion without challenging the observation.

Name the counterfactual out loud

03

What would have changed anyway? Order volumes, headcount, seasonality, a system upgrade landing the same quarter. A saving that coincides with a quiet period is not a saving, and the person who spots that after you present is not going to be gentle about it.

Pick the number before go-live and write it down

04

One primary metric, agreed and recorded before anything is switched on. Metrics chosen afterwards get chosen to flatter the project — by everyone, unconsciously, including you. Writing it down in advance is the only defence.

Fix the measurement date at kickoff

05

Not 'when it settles down', which means 'when it looks good'. The study's bar is six months; for a narrow automation six to twelve weeks usually shows the shape. Put the date in the diary before you start.

Record the full cost, including internal time

06

Model, hosting, build days — and the internal hours your own people spend on it, which is the line everyone forgets and often the largest single item. A return calculated against the invoice only is not a return.

Exhibit 2 — The baseline sheet

Six things to write down before anyone switches anything on

RecordHow to capture itWorked example
The taskOne named process, narrowly defined. Not “admin”.Re-keying job sheets into the ERP
Who does itRole and headcount, because the conversion rate differs by role.2 × scheduling coordinator
How long, nowCounted over two weeks. Not estimated from memory.6.0 hrs each per week
How oftenVolume per week, so you can separate a saving from a quiet quarter.~180 job sheets / week
The one metricChosen and written down before go-live.Hours/week on re-keying
The measurement dateFixed at kickoff, in the diary. Not “when it settles down”.12 weeks after go-live

Author's own method, not a published standard. The worked example is illustrative. Convert hours to money using the Admin Tax Index multipliers, which are derived from ONS ASHE 2025 provisional Table 14 median pay and paid hours, plus employer NIC and minimum auto-enrolment pension.

Turning hours into a number you can defend

Step two needs an instrument, and this is where the Admin Tax Index earns its keep. It prices one hour a week of admin, per person, per year, by role — derived from ONS ASHE median pay and each occupation's own median paid hours, with employer NIC and minimum pension added.

That gives you a published, checkable multiplier instead of a number someone made up in a meeting. If you free six hours a week from a production manager, the index says roughly £10,900 a year. Two hours a week from each of twelve drivers is roughly £23,100. Both figures are floors — they exclude overheads, vehicles, equipment and cover — which is exactly what you want when presenting to a sceptic. A conservative number that survives scrutiny beats an ambitious one that does not.

Then set that against the full cost. The cost benchmark gives you the split: on a typical build the model itself is around 1.4% of year-one spend, and delivery is nearly all of it.

Now you have a before, an after, a conversion someone can check, and a cost. That is what condition two and condition three were asking for.

The uncomfortable part

If you run this properly, you have to accept the answer.

Sometimes the honest measurement shows a project that did not pay back — the hours came out lower than expected, or the process changed underneath the automation, or the counterfactual ate the gain. That is a real result and it is worth having. It is much cheaper to learn it at twelve weeks on one process than at two years across five.

It also tends to reveal something more useful than a yes or no. In most of the cases I have measured, the variance is not in the technology at all. It is in whether the process was redesigned around the tool or the tool was bolted onto the existing process, which is the pattern behind most AI project failure and the reason the 5% look so different from the 95%.

Where to start

Pick the process you would automate first. Spend two weeks counting how long it currently takes and how often it happens, before you talk to a supplier. That single act moves you out of the 95% as the study defines it, and it costs you nothing but attention.

If you would like help choosing which process to point that at, the hidden waste audit is built for exactly this, or book a call and we will pick one together.

Frequently Asked Questions

Do 95% of AI pilots really fail?

The figure comes from MIT Project NANDA's "The GenAI Divide: State of AI in Business 2025", published July 2025, based on 300+ publicly disclosed initiatives, interviews with 52 organisations and 153 senior leaders surveyed at four conferences. Its success criterion was deployment beyond pilot with measurable KPIs and ROI impact measured six months post-pilot. That definition means a project delivering genuine benefit without a recorded baseline is counted as a failure, so the statistic reflects measurement discipline as much as technology outcomes. The report also states its sample may not represent all segments and that selection bias is possible.

What did the MIT GenAI Divide report define as success?

Deployment beyond the pilot phase with measurable KPIs, with ROI impact measured six months post-pilot, adjusted for department size. All three conditions had to hold. The requirement for measurable KPIs and for ROI to be actually measured means that organisations which did not capture a before-state could not satisfy the definition regardless of how well the technology performed.

How do you measure ROI on an AI project?

Capture the baseline before you buy, because the before-state is unrecoverable afterwards. Count in hours first and convert to money second, keeping the two separate so the conversion can be challenged independently. Name the counterfactual — what would have changed anyway through volume, headcount or seasonality. Agree one primary metric before go-live and record it. Fix the measurement date at kickoff rather than waiting until results look favourable. And record the full cost including internal staff time, which is usually the largest forgotten item.

How long should you wait before measuring an AI project's return?

The MIT study used six months post-pilot. For a narrow automation such as document extraction, six to twelve weeks is usually enough to see the shape of the result. The important thing is to fix the date at kickoff rather than measuring when the numbers look good, because a date chosen after the fact will be chosen to flatter the project.

Is it better to build AI in-house or buy from a vendor?

The MIT report found specialised vendor solutions succeeded at around 67% against roughly 33% for internal builds, and that back-office automation outperformed sales and marketing deployments. That is a meaningful signal in favour of buying where a suitable product exists. Building is easier to justify where the process is genuinely specific to your operation and no product fits it, but the default assumption should be that a proven product outperforms a first internal attempt.


Sources. The GenAI Divide: State of AI in Business 2025, MIT Project NANDA, July 2025 — methodology, success criterion and stated limitations via aigl.blog's summary, corroborated by Fortune, 18 August 2025, The Register, 18 August 2025 and the Marketing AI Institute critique. The report PDF at mlq.ai returned HTTP 403 to automated fetch on 29 July 2026, so no claim here rests on my direct reading of the primary document. Vendor-versus-build success rates per Fortune's account. Hours-to-cost conversion method and multipliers from the Admin Tax Index 2026, built on ONS ASHE 2025 provisional Table 14. All URLs fetched and verified 29 July 2026.