Auxeon
Menu

Auxeon / Navigation

← All insights

Why most AI pilots never reach the P&L

The research on stalled pilots points to the models and the work around them.

Share on LinkedIn ↗

The most cited AI statistic of the past year is a discouraging one. In The GenAI Divide: State of AI in Business 2025, researchers at MIT's Project NANDA reported that about 95% of organizations studied reported no measurable return on their generative AI investments, despite an estimated $30 to $40 billion in enterprise spending.1

The figure deserves a careful reading. It does not say the models failed, or that the pilots produced nothing anyone valued. It says most organizations studied reported no measurable financial return. That is a narrower finding, and a more useful one, because it describes a problem organizations can fix.

What the research points to

The report's authors located the cause in what they called a learning gap: tools that could not retain feedback, adapt to context, or improve over time, deployed into organizations that had not changed how the work itself moved.1 General-purpose assistants were widely adopted for individual tasks. Far fewer custom systems reached production inside real workflows.

Gartner's outlook for agentic AI points the same way. In June 2025, the firm predicted that more than 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls.2 The risks involve the models and the design of the work around them.

What separates the pilots that ship

Across the research and our own work, the pilots that reach the P&L share five traits.

  • They start with a measured baseline. If no one recorded how long qualification took before the pilot, no one can show it improved afterward. The measure has to exist before the build.
  • They sit inside a workflow, not beside it. A tool employees must remember to open is a tool they will stop opening. Systems that act where the work already happens, in the CRM, the inbox, or the intake form, get used.
  • They have an owner. Someone is accountable for the outcome, not just the software. When results drift, that person notices and acts.
  • They are scoped to a decision, not a department. "Use AI in sales" fails. "Respond to every inbound inquiry within five minutes with a qualified next step" can be built, tested, and measured.
  • They include a human gate. Consequential steps route to a named person with the context to decide. Teams trust systems that know their limits.

The uncomfortable implication

Most of this work happens before any model is chosen: mapping how work moves today, deciding where intelligence belongs, and agreeing on what success means. It is less exciting than a demonstration, which is why so many organizations skip it. The research suggests they pay for that later.

It is also why Auxeon begins every engagement with a Systems Review rather than a tool. The map comes first because it decides whether anything built afterward will show up in the numbers.

Sources

  1. Aditya Challapally et al., The GenAI Divide: State of AI in Business 2025, MIT Project NANDA, July 2025. As reported in "MIT Report Finds Most AI Business Investments Fail, Reveals 'GenAI Divide'," Virtualization Review, August 19, 2025.↩1↩2
  2. "Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027," Gartner, June 25, 2025.↩
Related capabilityManaged Operations

THE STARTING POINT

Establish the sequence before the spend.

Set the priorities, the architecture, and what comes first.