AI · Automation · Buying decision

95% of AI pilots never reach production: what the other 5% do differently

The demo went beautifully. The committee applauded. Six months later nobody is quite sure what happened to it. That is, by a distance, the most common ending for an AI pilot — and the data behind it is no longer a matter of opinion.

This guide covers what MIT and Gartner found, why pilots die, what separates a demo from a live process, how to detect agent washing, and the seven filters worth applying before approving the next one.

A laboratory tray of dimmed, cracked AI pilots beside a single bright orb crossing a glass gateway and docking into a running process lane with an amber checkpoint
AR
Owner of Dokuflex
Updated: 10 September 2026

For general management, operations and IT. An evaluation guide for anyone who has to approve, extend or close an artificial intelligence pilot. Includes vendor selection criteria and the metrics to fix before you start.

Direct answer

AI pilots do not die because of the model: they die because they never touched the real process. MIT's study found no measurable return in 95% of the organisations analysed, and the pattern shared by the remaining 5% is always the same: a bounded case, a baseline measured before starting, AI running inside the workflow, a defined path for exceptions, and a business owner with a budget.

The numbers: what MIT and Gartner found

The most repeated headline of the past year comes from The GenAI Divide: State of AI in Business, by MIT NANDA, built on more than three hundred reviewed initiatives, some fifty interviews and around a hundred and fifty executive survey responses. Its conclusion: 95% of organisations obtained no measurable return from their generative AI pilots.

It deserves a precise reading, because the nuance is everything. It does not say the tools do not work. It says they did not translate into a visible number in the accounts. Most stopped at individual productivity gains — someone writing their email faster shows up in no indicator — without ever reaching the process that invoices, collects or delivers.

Gartner arrives at the same place from another angle. In June 2025 it predicted that over 40% of agentic AI projects will be cancelled before the end of 2027, because of escalating costs, unclear business value or inadequate risk controls. Two months later it estimated that 40% of enterprise applications will feature task-specific AI agents by the end of 2026, up from less than 5% in 2025.

Both figures together

AI will be everywhere and most projects that pursue it as a project will be cancelled. That is not a contradiction: it means AI will arrive embedded in processes and applications that already work, not as a standalone initiative with its own steering committee. Treating it as a separate programme is the expensive mistake.

The five reasons a pilot dies

None of them is about model quality. All of them are about the environment the model was asked to work in.

  1. It lived beside the process, not inside it. The pilot was a separate screen you had to upload documents into and copy results out of. That works with twenty test cases and is abandoned at two hundred, because the shuffling eats the saving.
  2. Nobody decided what happens to the odd ones. AI gets the standard case right; the value sits in what happens to the 8% that is not standard. Without an explicit rule — who it goes to, at what priority, within what deadline — the exception returns to email and the process stops being a process.
  3. There was no baseline. If nobody measured how long the process took and what it cost before, the improvement cannot be proven. And what cannot be proven does not get budgeted next year. This is the cheapest failure to avoid and the most frequent.
  4. The owner was technical. Pilots that survive are requested by whoever suffers the problem — the head of admin, of purchasing, of people — because that is who defends the budget when a decision is due. A systems sponsor can build it, but cannot justify it.
  5. Cost per case was never calculated at scale. Model consumption, human review, retries and maintenance are irrelevant at a hundred documents a month and decisive at a hundred thousand. That calculation belongs in the ROI calculator before you sign anything, not after.

MIT's research adds an underlying cause that fits all five: deployed systems do not learn from use. They do not retain the reviewer's judgement, do not adapt to the company's context and do not improve over time. They start at a reasonable hit rate and stay there, while the team was expecting a rising curve.

What separates a demo from a live process

The distance between the two columns below is where the 95% is lost. It is worth reviewing before the demo, not after.

Dimension In the demo In production
The data Twenty hand-picked, clean, legible examples. Skewed scans, faxes, phone photos and one supplier's peculiar format.
The intake Someone uploads the file by hand. It arrives on its own by email, portal or integration, and is identified and routed without intervention.
The error It is noted and you move to the next example. It has a path: detected, escalated to a person, closed within a deadline.
The trail Not needed. Who decided what, with which version and on what evidence. This is what you show in an audit.
Permissions One user who sees everything. Each user sees their own; the assistant inherits those permissions and cannot widen them.
The cost Irrelevant. Cents per case multiplied by annual volume, plus human review.

Not one of those six rows is fixed by a better model. All of them are fixed by a process around the model: routing, exceptions, permissions, logging and measurement. That is exactly what a BPM does, and it is why the AI that pays off is usually built on top of one.

Agent washing: agent or rebranded chatbot?

Gartner named the phenomenon: agent washing, the practice of relabelling existing products — assistants, RPA robots, rules-based chatbots — as AI agents. Its estimate, in the same note that forecasts the cancellations, is blunt: out of thousands of vendors presenting themselves as agentic, only around 130 genuinely are.

Four questions are enough to separate substance from packaging in a sales meeting:

  • Does it decide, or only answer? An agent chooses between possible paths based on the state of the case. If the route is fixed in advance, it is a workflow with natural language on top — perfectly useful, but not the same thing and it should not cost the same.
  • Does it act on a real system? Writing to the ERP, posting the entry, opening the case. If it only returns text for a person to copy, the copying keeps the saving.
  • Does it record why it did what it did? With no decision log there is no audit, and with no audit there is no deployment in a regulated process.
  • Where is the brake? What it can never do, which amount triggers human validation, who revokes its permissions. A vague answer means risk control does not exist — one of the three cancellation causes Gartner cites.

We set out the wider framework in what BPM low-code with AI delivers and the practical implementation path in how to implement BPM low-code with AI.

Buy or build: the uncomfortable finding

It is the decision that consumes most committee time and the one MIT's study answers with least ambiguity: deployments backed by specialist vendors succeeded roughly twice as often as internal builds.

This is not about talent. It is about scope. When a company decides to build, it is not building «a call to a model»; it is also taking on continuous evaluation, version management, access control, traceability, exception handling and migration when the model it relies on becomes obsolete in eighteen months. None of that appears in the initial estimate, and all of it is what eventually consumes the team.

The working rule: build what differentiates you, buy what costs you. A proprietary risk model built on data only you hold can justify the effort. Classifying supplier invoices, extracting fields from a delivery note or routing incoming post does not: that is already a product feature, and it costs less than the first month of the build.

Seven filters before approving the next pilot

If a pilot does not clear all seven, the risk is not that it will go badly: it is that you will not be able to prove it went well. Which is worse.

  1. A named business owner. The person who suffers the problem today and will defend the budget tomorrow.
  2. A measured baseline. Cases per month, minutes per case, cost per case and current error rate. Before anything is touched.
  3. A success criterion agreed in writing. One number and one date: «cut case closure from 9 days to 3, by 30 November».
  4. Real, ugly data. Last quarter's cases exactly as they arrived, including the three suppliers who send impossible invoice formats.
  5. A defined path for exceptions. Who they go to, with what deadline, at what priority. Without this there is no production.
  6. Integration agreed on day one. Where the case comes in and where the result goes out. If the answer is «we will look at that in phase two», phase two never comes.
  7. A closing date. Four to eight weeks, with a binary decision: production with a budget, or closed with the lessons written down.

The seventh filter is the most uncomfortable and the most valuable. A pilot with no closing date never fails: it lingers indefinitely, consuming attention. Which is precisely what that 95% describes.

How Dokuflex solves it: AI as a process step, not a project

Our position is the one the data supports: the problem is almost never the model, it is everything around it. That is why in Dokuflex BPM low-code AI is not a separate product, it is a type of task inside the flow.

  • The case arrives on its own. Email, portal, integration or capture from document management. Nobody is uploading files into a separate screen, which is the first cause of abandonment.
  • The exception already has a path. If confidence falls below the threshold or the amount exceeds the limit, the case is escalated to the right person with a deadline. Routing is part of the flow, not a pending task.
  • Measurement is built in. Every step is logged with a timestamp, so the before-and-after comparison does not have to be assembled: it is read off. It is the same trail that process mining exploits.
  • Changing means configuring. Adjusting a threshold, adding a validation or changing who handles an exception means editing the process and publishing the version. That short cycle is what gets you to week eight with results rather than a status report. We cover it in how to cut BPM implementation time.

And intelligent document processing handles the least glamorous row in that table: the genuinely ugly documents, the ones that sink pilots built on twenty perfect PDFs.

Book a demo with your own documents, not ours →

How to choose the first case

Instinct says pick the most impressive process. Resist it: the first case is not there to impress, it is there to prove the full cycle works in your organisation. Four conditions:

  • High, repetitive volume. Without volume there is no visible saving, however good the technology.
  • Clear rules. If two internal experts cannot agree on what «correct» means, the problem is not an AI problem.
  • Tolerable, reversible error. A badly extracted field corrected at review, not a decision that reaches the customer unfiltered.
  • Available history. It is what lets you compare against the baseline and end the argument with data.

Supplier invoice capture, sorting incoming post, document checks on a new customer. Boring cases, and that is why they win. If you want to price it before you start, how RPA and BPM fit together helps order the debate between robots, workflows and AI.

Frequently asked questions

Is it true that 95% of AI pilots fail? +

The figure comes from The GenAI Divide: State of AI in Business 2025, by MIT NANDA, and it is worth reading precisely: it does not say that 95% of pilots fail technically, it says that 95% of organisations obtained no measurable return in their profit and loss. The distinction matters, because the failure is not in the model but in the jump from experiment to the process that invoices, collects or delivers.

Why does a pilot that demoed well end up dying? +

For five recurring reasons: the pilot lived beside the process rather than inside it, so somebody had to copy and paste; nobody defined what happens to exceptions; there was no baseline figure, so the improvement could not be proven; there was no business owner, only a technical sponsor; and the cost per case, irrelevant at a hundred documents, stopped being irrelevant at a hundred thousand.

What is agent washing? +

It is the practice of relabelling an existing product as an «AI agent»: a chat assistant, an RPA robot or a rules-based chatbot. Gartner named it in June 2025 while warning that over 40% of agentic AI projects will be cancelled before the end of 2027, and estimated that out of thousands of vendors claiming to be agentic only around 130 genuinely are. The practical test: if it cannot choose between several paths, act on a real system and leave a record of why it did so, it is not an agent.

Should you buy an AI solution or build it in house? +

MIT's research points in one direction: deployments backed by specialist vendors succeeded roughly twice as often as internal builds. The reason is not the quality of your team, it is scope: building means also taking on evaluation, supervision, versioning, traceability and maintenance when the model changes. For a genuinely differentiating use case that can pay off; for classifying invoices, it does not.

How long should an AI pilot last? +

Four to eight weeks. If it takes longer it is almost always because it has become an integration project disguised as a trial. A well-framed pilot starts from a baseline measured before anything is touched, runs on real cases from the last quarter and ends with a binary decision: it goes to production with an owner and a budget, or it is closed.

What should you measure to know whether an AI pilot worked? +

Four figures, none of which is model accuracy: end-to-end case time against the baseline; the share of cases completed with no human intervention; the error rate that reaches the end customer; and cost per case, including model consumption, human review and maintenance. Accuracy is an internal indicator; those four are what a board understands.

Where should you start with AI in processes? +

With a high-volume process that has clear rules, repetitive input documents and a tolerable, reversible cost of error: supplier invoice capture, sorting incoming correspondence, document checks on a new customer. They are boring cases, and that is exactly why they work: there is history to compare against, a prior figure to beat, and a mistake is caught and corrected without serious consequences.

Sources

Next step

Start with the number, not the demo

Work out in 60 seconds what the process you want to automate costs you today. No email, no sales call. With that figure on the table, every conversation afterwards — with us or anyone else — gets a lot shorter.