The question arrives at a predictable point in every enterprise AI deployment, somewhere past the pilot, somewhere before the renewal, and it usually arrives from finance. Is it worth it? The licences have a number attached. The tokens have a number attached. The question assumes the value has a number attached too, and that someone has simply neglected to write it down. The cost side of the ledger is the easy half. The value side is where the trouble starts, and the trouble has a name: attribution. Separating what the tool contributed from a good quarter, a strong new hire, or a team that would have improved anyway is the entire problem, and most organisations discover it only after they have promised an answer.
Medicine solved this problem a century ago, and the solution was the control group. You do not learn whether a treatment works by giving it to everyone and asking how they feel. You learn by withholding it from some, randomly, and measuring the difference. A healthcare company will not ship a clinical protocol without trial evidence. The same company will roll a productivity tool out to its entire workforce, celebrate the adoption curve, and then ask, with no apparent sense of contradiction, whether the tool is working.
There is a cruel arithmetic here. The better an adoption programme performs, the faster it destroys the evidence of its own worth. At ninety-five per cent adoption there is no one left to compare against. The moment the question can finally be asked with confidence (everyone is using it, surely now we can measure) is precisely the moment it can no longer be answered with rigour. The organisations that want a defensible number have to carve out their holdouts early, while adoption is still climbing, which is exactly when nobody wants to think about measurement. Success eats the control group.
The second trap sits on the other side of the ledger. The value of any tool is a rate: output per unit of input. And giving a person an AI assistant raises both. The output is visible: the draft produced, the code written, the analysis delivered. What rises invisibly alongside it is the input: the reading, the checking, the quiet clawing back of changes that were confidently wrong. Call it the verification tax. It does not appear on any dashboard, and the people paying it are the least reliable witnesses to its size. In one widely cited study, experienced developers estimated that AI assistance had made them about twenty per cent faster; when their work was actually measured, it had made them slower. They were not lying. They were confidently wrong, which is worse, because confident error is what self-reported productivity surveys are made of.
This is the context for a pair of findings that ought to be more famous than they are. One research group at MIT found that roughly ninety-five per cent of enterprise AI pilots produced no measurable impact on the profit-and-loss statement. McKinsey found that while the overwhelming majority of firms report active AI usage, only around six per cent report a meaningful contribution to earnings. Read carelessly, these numbers say the technology is failing. Read carefully, they say something more specific: the gains are real at the level of individual workflows (a support team’s handle time, an engineering team’s cycle time), and they evaporate somewhere on the way to the income statement. The failure is mostly one of measurement and integration, not of the tools. The honest posture, and the rare one, is to measure deeply where the instruments exist and decline to promise what they cannot show.
There is a second uncomfortable finding underneath the first. When the consultancies decompose where enterprise AI value actually comes from, the model itself accounts for roughly ten per cent. The rest is the apparatus around it: the data plumbing, the integrations, the redesign of the work itself. This matches what the benchmarks have been saying for a while: output quality on everyday knowledge tasks reached good enough a few model generations ago, and the improvements now concentrate in specialised domains, in long autonomous loops and competition mathematics and multi-file software changes. One widely shared head-to-head this month found the newest frontier model producing the same everyday result as its cheaper rivals at nine times the cost and twenty times the wait. The implication for anyone building an ROI story is clarifying. The infrastructure compounds; the model depreciates. A value case built on the harness does not reset every time a lab ships a new release. A value case built on the model resets constantly, which is to say it is not a case at all.
Part of the difficulty is that “is it worth it” is three questions wearing one coat. The first is justification (did the spend return more than it cost), and it belongs to finance, answered with displaced spend and instrumented workflows. The second is allocation (where should the next dollar of effort go), and it belongs to whoever runs the programme, answered by the consistent finding that gains concentrate among novices and thin out among experts. The third is option value: the capabilities that did not exist before. An agent that investigates incidents overnight is not faster than the old way; there was no old way. Forcing a dollar figure onto option value produces numbers that embarrass everyone who reads them. The right treatment is to narrate it honestly alongside the measured results, and to resist the spreadsheet’s gravitational pull toward false precision.
None of this requires a statistics department. It requires two questions, asked reflexively, every time a return-on-investment claim crosses the desk: a vendor’s, a consultant’s, your own. What is the counterfactual? And is the input held honest? Most claims survive neither. The few that do are the only ones worth building a strategy on.