The 95% AI Pilot Failure Rate Is Good News

Why the MIT number is a portfolio management finding, not a technology verdict, and the three governance mechanisms that separate an AI experiment pipeline from a capital leak.

By Rob Purks, Founder and Operating Partner, Lumerai Advisors

Key Takeaways

  • A 95% AI pilot failure rate is meaningless without two companion metrics: average cost per terminated initiative and median time from funding to termination.
  • Pharmaceutical development fails on roughly 90% of compounds entering clinical trials. That is a functioning pipeline, not a broken industry, because failure is designed to arrive early and cheap.
  • Most enterprises do not have an AI pilot problem. They have a termination problem: initiatives are deprioritized quietly rather than stopped deliberately.
  • Termination Debt is the accumulated liability of commitments an organization has no mechanism to retire. It is the companion liability to Governance Debt.
  • Three mechanisms convert failure into value: pre-registered kill criteria, time-boxed experiments with default expiry, and learning velocity as the governing metric.

When MIT’s Project NANDA reported that roughly 95% of enterprise generative AI pilots had produced no measurable financial return, the number moved markets. It was read as a verdict on the technology.

It is not a verdict on the technology. It is a description of an experiment portfolio, and on its own it tells you almost nothing about whether the money was well spent.

What the MIT AI Pilot Failure Rate Actually Measures

Consider pharmaceutical development. A ten-year cross-industry study by the Biotechnology Innovation Organization, analyzing nearly 10,000 clinical trial phase transitions, found a cumulative approval rate of 9.6 percent – meaning roughly nine out of ten compounds entering clinical trials never reach approval. Nobody reads that number and concludes drug discovery is broken. The industry built its entire operating model around the expectation of failure: phased trials, endpoints defined before enrollment, and stage gates that terminate a compound the moment it misses. Failure is not the flaw in that system. Failure is the system working, cheaply and early.

The question is not what percentage of your AI pilots failed. It is what each failure cost you, and how long it took you to find out.

A 95% failure rate on experiments costing $40,000 each and concluding in six weeks is a healthy innovation portfolio. A 95% failure rate on initiatives costing $4 million each and drifting for three years is a capital allocation crisis. The headline figure is identical. The underlying businesses are not remotely comparable.

The Two Metrics Missing From Every Enterprise AI Portfolio Review

Almost no organization I encounter can produce the two numbers that would tell them which of those businesses they are:

  • Cost per termination. Average fully loaded spend on initiatives that reached a documented stop decision.
  • Time to termination. Median elapsed time from funding release to that decision.

These are the metrics that separate a research pipeline from a slow leak. They are largely unmeasured because most enterprises have no discrete moment they can point to and call termination.

Termination Debt: Why Enterprise AI Initiatives Never Actually Die

In practice, enterprise AI initiatives rarely get killed. They get quietly deprioritized. Headcount thins. The steering committee moves from monthly to quarterly to dormant. Eighteen months later the initiative surfaces on a portfolio review as amber, and nobody in the room can say when it stopped being real.

I call this Termination Debt: the accumulated liability of commitments an organization has no mechanism to retire. It is the companion to Governance Debt, the concept I introduced in earlier work on AI ROI. Governance Debt is what you owe because you deployed faster than you could govern. Termination Debt is what you owe because you committed faster than you could stop.

It is expensive in ways that never appear in a business case. Direct spend is the smallest component. The larger costs are senior attention consumed by work already known to be going nowhere, and delivery capacity locked up in initiatives nobody will defend.

Credibility erodes too, each time an executive team announces something that later evaporates without explanation. That erosion compounds. By the third undead pilot, the organization has learned that leadership announcements do not predict outcomes, and the next initiative, including the good one, starts with a deficit of belief. This is governance debt in its most literal form: commitments accumulated faster than any mechanism built to retire them.

Three Mechanisms That Convert Failure Into an Asset

The distinction between a portfolio that learns and one that leaks is not talent or technology. It is whether termination is designed in advance or negotiated after the fact. Three mechanisms do most of the work.

  • Pre-registered kill criteria. Borrowed directly from clinical trial reform. Before funding is released, the sponsor writes down what result would cause the initiative to stop, and that document is filed with the funding decision rather than with the project team. Pre-registration matters because it removes the negotiation. When the criteria are written after results arrive, they are always written to accommodate the results. The reason to pre-register is not scientific rigor for its own sake. It is that a team facing a threshold it agreed to twelve weeks ago is in a fundamentally different conversation than one defending its existence.
  • Time-boxed experiments with default expiry. Every experiment receives a hard expiry, not a target date. At expiry it converts to production, receives explicit refunding against new criteria, or ends. Silence defaults to ending. This inverts the usual pattern, in which continuation is the default and termination requires someone to expend political capital.
  • Learning velocity as the governing metric. Most portfolios report pilots launched, which measures activity and creates a perverse incentive to start more than the organization can conclude. The better metric is how many validated answers to consequential questions the portfolio produced this quarter, counting negative results. A pilot that establishes cheaply and definitively that a use case will not work has delivered value. It has removed a plausible option from the roadmap and returned the capital that would have chased it.

What Does This Discipline Look Like in Practice?

Pre-registered kill criteria and time-boxed experiments recreate deliberately what a hard deadline once forced for free. I once led the IT technology build for a multibillion-dollar wireless startup in Latin America – a global team spread across four continents, standing up billing, CRM, and the full operational stack from zero against a hard twelve-month launch deadline. 

Five weeks into requirements and design for the prepaid billing platform, a Fortune 100 vendor we had selected in the RFP process showed gaps serious enough to put the launch at risk – a critical exposure in a market where prepaid is the dominant form of wireless service. 

We did not spend months negotiating a remediation plan. We shifted to the second-ranked vendor from the original evaluation, absorbed the rework in requirements and solution design, and kept the launch date intact. The switch cost weeks. Riding out the first vendor’s gaps in the hope they would resolve could have cost the launch itself. 

That is the same trade pre-registered kill criteria are built to force: a threshold agreed to in advance, acted on the moment it was crossed, instead of negotiated after the fact once the damage was already done.

So What Should Change After Reading This?

The MIT number is not evidence that enterprise AI has failed. It is evidence that a very large number of organizations are running experiments without having built the apparatus that makes experimentation worthwhile.

The companies that pull ahead over the next three years will not be the ones with a higher pilot success rate. They will be the ones that failed at roughly the same rate for a fraction of the cost, learned faster from each result, and still had the capital and the credibility intact when the use case that actually mattered finally appeared.

A 95% failure rate is what a functioning pipeline looks like. The work is making failure cheap, fast, and safe to report.

Which means the 95% is not a reason to underwrite fewer AI-driven gains. It is a reason to underwrite them differently – to fund a portfolio of cheap, fast, well-governed experiments instead of a small number of expensive, slow, ungoverned ones, and to treat the discipline of killing things as a diligence-grade capability rather than a cultural nicety.

The portfolio companies that win the next cycle will not be the ones that avoided the 95%. They will be the ones that got through it in ninety-day increments while everyone else was still deciding whether to admit it.

Related reading: Governance Debt Is Already Eating Your ROI – why the AI liability standard diligence cannot see is already priced into your return. And The AI ROI Panic Is About to Create the Next Legacy System – where I first defined Governance Debt.

Frequently Asked Questions

Why do 95% of enterprise AI pilots fail?

MIT’s 2025 NANDA research attributed the failures not to model quality but to conditions set before the model was built: fragmented and ungoverned production data, no named owner after deployment, and workflows that were never redesigned to act on AI output. The failure rate reflects organizational readiness, not the technology.

Is a 95% AI pilot failure rate normal?

For genuine experimentation, yes. High failure rates are the expected shape of any portfolio where returns come from a small number of large wins. The rate becomes a problem only when individual failures are expensive and slow – when pilots run eighteen months and are never formally closed, they stop producing learning and start producing Governance Debt.

How should Operating Partners measure AI programs?

Replace success rate with learning per quarter: the number of decision-grade answers the organization purchased in the last ninety days and what each one cost. Success rate incentivizes teams to keep failed pilots alive; learning per quarter incentivizes them to resolve questions quickly.

What are pre-registered kill criteria?

A written, signed threshold and date, agreed before a pilot begins, that defines the result which ends it. Pre-registration matters because criteria set after the data arrives become a negotiation rather than a decision rule.

What should a board ask about AI pilots?

How many experiments were started versus formally closed last quarter, what the written kill criterion is for each active pilot and who signs it, what the organization learned that changed a decision, and the average time from pilot start to formal go/no-go. These assess decision-making capability rather than technical merit.

The Hidden Risk of AI: Building Transformation Programs for a Future That May Not Exist

The AI ROI Panic Is About to Create the Next Legacy System

Lumerai Technology Value Realization – Part 2