• Report
  • 6 min read

Where AI pilots stall on the way to production

In MIT NANDA's review, a quarter of enterprise AI pilots reached production. We followed a pilot through the stages where the 2025 and 2026 studies say it loses ground.

ResearchBy OliviaPublished 3 September 20266 min read8 sources

Pilots get the best conditions a company can offer. The team volunteered, the data was chosen for them and the sponsor wants to report a success. Production removes all three, and the research published since early 2025 keeps pointing at the same gap.

MIT NANDA’s review is the most quoted source on that gap. It covered more than 300 publicly disclosed AI initiatives, 52 structured interviews and 153 survey responses gathered between January and June 2025.1 Of the organisations that evaluated enterprise-grade AI tools, a third reached a pilot and a quarter of those pilots reached production. The sections below take the stages in the order a project meets them, and draw on Gartner, IBM, BCG and McKinsey where they add something MIT does not.

Enterprise-grade AI tools: from evaluation to production

Share of organisations reaching each stage, January to June 2025

Evaluated a tool60%
Reached a pilot20%
Reached production5%

Source: MIT NANDA, The GenAI Divide, 2025, as reported by Virtualization Review, 19 August 2025.1

Before the pilotStaff are already using AI, on accounts the company does not manage

Only 40% of companies in the MIT study had bought an official large language model subscription. At more than 90% of them, staff said they used personal AI tools for work.1

Two conclusions follow, and they pull in different directions. Demand already exists, because employees have found their own uses for these tools. The approved tool, when it arrives, competes with whatever people do on their phones, and it loses if it cannot handle the same tasks.

There is a security side as well. Work material is moving through accounts nobody in IT can see. A ban without a better approved option tends to push that use further out of sight, which is the opposite of what the ban was for.

Inside the pilotThe sample was cleaned by hand

A pilot team can fix its data manually. Nobody can do that across a whole business. Gartner surveyed 1,203 data management leaders and found that 63% of organisations either lack the data management practices AI needs or cannot say whether they have them. It expects organisations to abandon 60% of AI projects through 2026 where the data is not AI-ready.2

Do organisations have the data practices AI needs?

Share of data management leaders, Gartner

Lack them, or unsure63%
Report having themCalculated as 100 minus 6337%

Source: Gartner, 26 February 2025. The second bar is derived from the first.2

Other surveys say the same thing in their own terms. In Informatica’s 2026 study of 600 data leaders, 57% named data reliability as a key barrier to moving AI from pilot to production.3 Fivetran found that 42% of enterprises said more than half of their AI projects had been delayed, had underperformed or had failed because of data readiness.4

Gartner’s advice runs to five steps: align data to specific use cases, work out the governance requirements AI adds, make metadata active, prepare pipelines for training and for production, and keep testing the data once it is live.2 The first real test of a pilot, in our reading, is whether its last-mile data preparation can be automated at all.

Inside the pilotCorrections that vanish

MIT names learning capability as the main barrier it found. The tools in its sample often failed to keep feedback, adapt to a team’s context or improve with use. One mid-market manufacturing COO quoted in the study said that LinkedIn makes it sound as though everything has changed, while day-to-day operations look much the same.1

This is easy to test during a pilot and rarely tested. Ask five users to correct the system on the same task over two weeks, then check how many corrections still hold at the end without being repeated. A figure close to zero means the pilot is a demonstration, and in production it would repeat the same mistakes at full volume.

The approval meetingRunning costs show up after the decision

Gartner gives three reasons for its forecast that more than 40% of agentic AI projects will be cancelled by the end of 2027: escalating costs, unclear business value and inadequate risk controls.5 In McKinsey’s 2026 survey, one respondent in five said their organisation was limiting AI use because of operating costs.6

Costs come late because pilots are small. Usage is low, vendors sometimes discount it, and the people checking outputs are the enthusiasts who volunteered. The full bill arrives when the system opens to everyone.

CostWhy the pilot hid itQuestion for the approval meeting
Cost per run at full useVolumes were small and often discountedWhat does a run cost at ten times pilot volume?
Data preparationThe team cleaned the sample by handWhich preparation steps are automated, and who owns them?
Review of wrong answersPilot users reported errors informallyWho reviews outputs once the system is live, and how quickly?
Audit and loggingA controlled test did not need themWhat would an auditor ask to see?

After approvalLarge organisations need longer, and finish stronger

MIT reports that mid-market firms took about 90 days from pilot to full implementation, while large enterprises took nine months or more.1 IBM’s survey of 2,000 CEOs points to one reason: half said rapid investment had left them with disconnected, piecemeal technology.7

Time from pilot to full rollout

MIT NANDA, 2025

~90 daysmid-market firms
9+ monthslarge enterprises

Source: MIT NANDA via Virtualization Review, 2025.1

McKinsey’s 2026 figures show the other half of the story. Among respondents from large organisations, 40% now say they are scaling AI agents, against 22% at smaller ones.6 Size slows the start and seems to help later, perhaps because the platform, approval and ownership work that large firms do early pays back once projects are ready to scale.

For a sponsor, the practical point is the timeline. Nine months is normal for a large enterprise in this evidence. A plan written for three will report a delay in month four while the project is on schedule.

What the survivors shareThey change the work before they widen access

McKinsey found that nearly three-quarters of its high performers had fundamentally redesigned workflows because of AI, up from 55% a year earlier. Among other respondents the figure was about a quarter.6 BCG reports that 33% of its "future-built" companies use AI agents and spend about 15% of their AI budget on them, against 12% of scalers and almost no laggards.8

Share of each BCG group using AI agents

BCG survey of 1,250 senior executives, 2025

Future-built33%
Scalers12%
LaggardsAlmost none

Source: Boston Consulting Group, 30 September 2025.8

Across these studies the order of events is consistent. The organisations that get through pick a process, give it a business owner, agree what a good result looks like and only then put the model inside it. Wider access comes last.

For your next reviewSix symptoms and where to look first

SymptomLikely causeFirst check
Good demo, no production dateNobody in the business owns the workflowName the owner and the process the system will change
People prefer their personal chat toolThe approved tool cannot do a task staff rely onCollect the ten tasks people take to personal tools
Accuracy drops as the sample widensThe pilot data was cleaned by handAutomate preparation and measure error on unseen data
Different answers to the same business questionTerms are defined differently across sourcesPublish one definition per term, with an owner
Sponsor reports slippage in month fourThe plan assumed mid-market speedReset against a nine-month reference for large firms
Budget approved, then frozenRunning costs were never modelled at full volumePrice the system at ten times pilot use before the next review

Each row rests on the evidence above, and each needs your own data before it means anything for a particular project.

Sources and method

Each figure links to its original publication. All sources were published in 2025 or 2026. The MIT NANDA report is cited through trade press coverage, which is named in the entry.

  1. MIT Report Finds Most AI Business Investments Fail, Reveals "GenAI Divide" (coverage of MIT NANDA, The GenAI Divide: State of AI in Business 2025)Virtualization Review, 19 August 2025
  2. Lack of AI-Ready Data Puts AI Projects at RiskGartner, 26 February 2025
  3. New Global CDO Report Reveals Data Governance and AI Literacy as Key Accelerators in AI Adoption (CDO Insights 2026)Informatica, 27 January 2026
  4. Fivetran Report Finds Nearly Half of Enterprise AI Projects Fail Due to Poor Data ReadinessFivetran, 13 May 2025
  5. Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027Gartner, 25 June 2025
  6. The state of AI in 2026: On the road to ROIMcKinsey & Company, 25 August 2026
  7. IBM Study: CEOs Double Down on AI While Navigating Enterprise HurdlesIBM, 6 May 2025
  8. AI Leaders Outpace Laggards with Double the Revenue Growth and 40% More Cost SavingsBoston Consulting Group, 30 September 2025

This report summarises published research and is not professional advice. The MIT NANDA findings rest on the study’s own sample.

BIAAS 2027, Amsterdam

Hear it first-hand from the people doing the work

Two days of keynotes, panels and 1-on-1 meetings with senior data, analytics and AI leaders, on 23-24 March 2027.