The Measurement Crisis Behind the AI ROI Debate
Earlier this year we were about to sign a six-figure annual contract for a business analytics tool at Recurly. We didn't. We replaced it with a little additional model spend, likely single-digit thousands a year, and got a stronger result, because the value now lands where our people already work, inside Claude Enterprise, and composes with the rest of our AI ecosystem instead of sitting in a separate destination.
The savings make a good headline. They are not the interesting part. The interesting part is that the thing we were about to buy was priced the way that category has always been priced, per seat and per license, and the thing that replaced it is priced like a task. The economic unit moved underneath us, and most enterprise ROI math has not caught up to it.
That gap is the real story behind the most-quoted statistic in the AI investment debate. MIT Project NANDA's The GenAI Divide: State of AI in Business 2025 found that 95% of enterprise GenAI pilots showed no measurable P&L impact, and only 5% of integrated systems produced significant value. The number has become the rhetorical anchor for every skeptic's case against the current cycle, which is exactly why it is worth slowing down on. A figure doing this much load-bearing work, cited far more often than it is read, usually conceals more than it reveals. Depending on how you read it, the 95% is either the most damning indictment of the technology or the most revealing indictment of our measurement frameworks.
I think it is mostly the second, and I think the cost of getting that wrong is larger than the debate acknowledges. NANDA's own diagnosis points the same way: it traces the failures not to model quality, infrastructure, or talent but to a learning gap: systems that don't retain feedback, adapt to context, or improve over time. That is a claim about how these systems behave, and behavior that drifts and adapts is exactly what conventional ROI math was never built to measure. The skeptics have annexed the headline number. The body of the report reads closer to a measurement critique than a verdict on the technology.
This is not a defense of failed pilots. Most enterprise AI deployments deserve to be flagged as value failures, and the discipline of refusing to count shelfware as transformation is the right discipline. But underneath the legitimate cases of bad deployment is a quieter problem: the analytical machinery enterprises are using to grade AI investments was built for a class of software that AI-native systems are not. We are running SaaS-era ROI math against a non-SaaS economic substrate and treating the resulting confusion as evidence about the technology rather than evidence about the framework.
That distinction matters because it determines what you do next. If the 95% is a value problem, you tighten procurement. If it is partly a measurement problem, you tighten procurement and rebuild the instrumentation. Most boards I sit with are doing the first and skipping the second.
What SaaS ROI math actually assumed
The conventional response to any "measurement is broken" claim is correct in principle: TCO, NPV, payback period, productivity per FTE. These frameworks are deliberately technology-agnostic, and they have held up reasonably well across prior software waves. Pick a pre-deployment snapshot, measure the delta over a defined window, discount for confounders, re-baseline as needed. Done.
That response is correct in principle and incomplete in practice, because the frameworks were not actually technology-agnostic. They quietly assumed four properties of the system being measured:
- Fixed marginal cost per unit of consumption (a seat, a license, a server)
- Deterministic output quality per workflow step
- A stable baseline workflow on the other side of the comparison
- A bounded attribution surface: the software did roughly one identifiable thing
SaaS satisfied all four. The cost was per seat, the output of a CRM record update was the same on Tuesday as it was on Friday, the workflow being replaced was at least nominally documentable, and the software's contribution to the outcome was bounded enough to defend in a board deck.
AI-native software violates all four. And when you violate the assumptions a framework was built on, the framework does not fail loudly. It produces clean-looking numbers that are quietly wrong.
Where the assumptions break
Cost is no longer per seat. An AI-native application is not just a model. It is a model surrounded by a context store, tool integrations, orchestration, state management, sandboxes, and evaluation infrastructure. Each of those is an independent cost and quality variable that did not exist in the SaaS P&L. The economic unit is closer to a task with a stochastic success rate and a variable inference cost than it is to a seat. The analytics-tool swap I opened with is one instance of this: a SaaS-era TCO model holding cost constant per user would have scored the six-figure license we did not buy and stayed blind to the task-priced system that replaced it, mismeasuring the cost side before it ever reaches the value side.
Output is non-deterministic and drifts. The system you measure in month one is not the system you measure in month six: the model gets updated, the context accumulates, the prompts evolve, and, for the subset of vendors that deliver on it, the system incorporates signal from prior runs. A pre/post comparison assumes a stable treatment, and I am not aware of an enterprise finance team with a standard methodology for grading a treatment that improves itself during the measurement window.
The baseline was never static either; we just pretended. This is the part most ROI debates skip. The SaaS workflow being used as the pre-deployment baseline was not actually a clean reference point. It embedded years of human workarounds, tribal SOPs, undocumented escalation paths, and shadow processes that quietly absorbed variability. When AI replaces a workflow, the cost of explicitly encoding what humans tacitly knew becomes visible, and the AI looks more expensive than the thing it replaced. In reality, the prior baseline was undercounted. The humans were finishing the recipe for free, off the books.
Attribution is split across at least three layers. BCG's widely cited claim that roughly 70% of AI implementation value or failure traces to people and process, 20% to tech and data, and 10% to algorithms is a directional finding, not a law of physics. But even taken loosely, it has a brutal implication for ROI work: measuring the model layer in isolation captures a small fraction of the causal story. Most ROI debates I read are arguing about model performance and treating it as if it explained the outcome. There is no established enterprise discipline for attributing returns across harness, process, and model simultaneously. Lean and Six Sigma can handle moving baselines, but they assume the treatment is the same treatment across the measurement window. None of the standard kits handle all three layers moving at once.
The supply side has the same problem in a mirror
I wrote earlier this year about the ARR measurement crisis among AI-native vendors: pilots being booked as recurring revenue, inference whales distorting unit economics, churn rates that no traditional SaaS company would recognize. At the time, I treated it as a vendor governance story. I think it is actually the same story as the buyer-side ROI crisis, viewed from the other side of the contract.
If the vendor cannot cleanly report what is recurring because usage is consumption-based, churn is volatile, and seats do not map to value, then the buyer cannot cleanly report savings against a per-seat baseline either. One side is trying to annualize stochastic consumption into a clean ARR number. The other side is trying to depreciate stochastic consumption against a clean productivity-per-FTE number. Both are forcing a variable, task-priced, drifting system into a fixed-price subscription frame. Both get clean-looking numbers. Neither set of numbers is doing what its readers think it is doing.
This is one measurement crisis with two faces, and the industry is treating them as separate problems.
What the consensus view gets right
I want to be careful here, because the consensus view is not wrong. It is incomplete.
It is right that re-baselining is a known discipline and most enterprises are not doing the unglamorous work of instrumenting the workflow before turning the AI on. It is right that consumption-based ROI math is not new. Snowflake and Twilio buyers have been doing variants of it for years, and most enterprises have simply not transferred the muscle to AI. It is right that a meaningful share of so-called AI-native tools are SaaS with an LLM bolted on, do not actually learn post-deployment, and should not get the analytical accommodation of an adaptive system. Treating every vendor's "we learn from your data" claim as real gives bad products cover they did not earn.
And it is right that some AI use cases do have legible ROI today. Call-center deflection, code-completion acceptance rates, document-processing throughput: these are bounded enough, deterministic enough, and have stable enough baselines that conventional measurement works fine. The measurement crisis is not universal. It bites hardest exactly where the value is supposed to be largest: cross-functional agentic workflows, knowledge work augmentation, anything where the system is supposed to learn.
So the honest position is not "ROI measurement is broken, lower the bar." The honest position is: ROI measurement works on roughly the same surface of problems it has always worked on, and the surface where AI is supposed to create the most value is exactly the surface where the existing frameworks were never designed to operate.
What to actually do
The temptation when the framework is wrong is to either abandon the discipline or wait for someone else to invent a new one. Neither is the right move for a CFO or board this year. Here is what I would do instead.
Separate cost-side and value-side measurement, and instrument them differently. Stop reporting AI ROI as a single number against a single baseline. On the cost side, track per-task inference cost, harness cost (context, tools, orchestration, evaluation), and the fully loaded cost of the humans still in the loop. On the value side, track task volume, task success rate, and the cost of failure modes. The two sides should be reconciled at the executive level but not collapsed into one number, because the units are not the same and pretending they are is how you get clean nonsense.
Make the prior baseline pay its full freight. Before you stand up the AI, document the workarounds, shadow processes, exception handling, and tribal knowledge that the existing workflow consumes. If you cannot enumerate them, you do not have a baseline. You have a number. The single biggest correction most enterprises could make to their AI ROI math is to stop comparing the new system against an idealized version of the old one.
Re-baseline on a fixed cadence, and disclose it. If the system genuinely learns, then a twelve-month pre/post comparison is measuring two different systems and calling it one. Quarterly re-baselining with explicit disclosure of what changed in the treatment is closer to honest. This is uncomfortable for vendors who want a single hero number and uncomfortable for buyers who want a single procurement decision. Do it anyway.
Stop measuring the model when the question is about the system. If most of the value lives in people and process, then your ROI instrumentation needs to measure people and process changes, not model benchmarks. The model layer is the easiest to measure and the least explanatory. Most AI ROI dashboards I have seen invert this: they are model-heavy and process-light, because the model data is what the vendor ships and the process data is what the enterprise has to build.
Refuse the "adaptive system" defense from non-adaptive vendors. Most tools sold as "AI-native" do not meaningfully learn after deployment. They should be measured with conventional SaaS ROI frameworks, because that is what they are. The measurement crisis is real, but it is real for a specific class of systems, and extending it to the whole category is how shelfware gets justified. The diligence question is concrete: show me, in production, what the system knows in month six that it did not know in month one. If the vendor cannot answer, measure them as SaaS.
The harder problem underneath
The reason the AI ROI debate feels intractable is not that one side is right and the other is wrong. It is that both sides are mostly arguing about the value layer using instrumentation built for a different economic substrate, and the substrate change is invisible inside the argument.
The buyers who say "we cannot measure it" are often correct about the system and wrong about the conclusion, which is not that measurement is impossible but that the existing instrumentation is mismatched. The skeptics who say "if you cannot measure it, it is not there" are often correct about the discipline and wrong about the diagnosis, which is not that the value is absent but that the framework is grading the wrong thing.
The work, for any executive trying to govern an AI portfolio in 2026, is to rebuild the instrumentation before drawing conclusions from it. That is unglamorous, it does not produce a hero number for the board deck, and it will surface uncomfortable facts about how much of the prior baseline was being held together by undercounted human labor. It is also the only way the next round of ROI debates will be about the technology rather than about the measurement.
The 95% number is doing real work in the industry right now. Some of that work is correct: a lot of deployments are not paying back, and the discipline of saying so out loud is healthy. Some of it is not: a finding produced by mismatched instrumentation is being read as a finding about the underlying systems. Sorting which is which is the actual job. It is harder than picking a side in the debate, and it is the only version of the conversation that produces a decision worth defending.