There is a ritual now, observed every few months with the regularity of a season. A frontier lab announces a date, the date arrives, and knowledge workers everywhere open a livestream at their desks to watch a chart. The chart shows bars, and the newest bar is taller than the previous bars. In the chat, people type the word “insane.” By the afternoon, the timeline has concluded that everything has changed. By Monday morning, for most of the people who watched, nothing has. The email still gets drafted, the meeting still gets summarised, the ticket still gets triaged, and the output is indistinguishable from the output of the model that was, seventy-two hours earlier, declared obsolete.
The belief underneath the ritual is that the release applies to you, that the taller bar represents some increment of capability that will arrive in your work the way a pay rise arrives in your account. It is a reasonable belief, encouraged by the launch post and the vendor deck alike, and for roughly eighty per cent of any organisation it is wrong. The evidence sits in plain view, one chart below the headline figure. On the evaluations that resemble ordinary knowledge work (comprehension, drafting, analysis, the great undifferentiated middle of the working day), the frontier models have saturated into a cluster so tight that distinguishing them requires instruments more sensitive than any actual task. MMLU-Pro, the benchmark built specifically because its predecessor had become too easy, is itself filling up. Meanwhile, on the benchmarks that measure the specialised work (AIME’s competition mathematics, SWE-bench’s multi-file software changes, the long autonomous loops where an agent works unsupervised for hours), the gaps between generations are widening. The frontier is still moving. It is just moving away from the average desk.
The misconception has a mechanism. Launch demonstrations are virtuoso pieces by design (the competition problem solved live, the agent running for six unsupervised hours), because virtuosity is what changed and virtuosity is what films well. The viewer watches a performance they will never give and concludes that their own work is about to improve.
A widely shared head-to-head this month made the shape of this visible. On everyday tasks, the newest frontier model produced the same result as its cheaper rivals at roughly nine times the cost and twenty times the wait, a premium paid, in effect, for the privilege of the logo. Then the same model was pointed at a targeted optimisation problem, and for forty dollars and two hours of unsupervised looping it took a layout resolver from microsecond to nanosecond scale, a result the engineer running the experiment conceded he could not have reached himself. Both findings are true. They are findings about different people. The engineers are not wrong to be excited. They are simply a minority who keep forgetting they are one.
For everyone else, the model was never the main event. When BCG decomposed where enterprise AI value actually comes from, the model accounted for about ten per cent; data and infrastructure for twenty; people and process for seventy. The other ninety per cent is unglamorous and entirely model-agnostic: the assistant knowing your products, your terminology, the reason the returns policy has that strange exception; the connection into the ticketing system and the finance stack, where the work actually lives, rather than a chat window adjacent to it; the redesign of the workflow itself. A city can upgrade its reservoir every quarter and issue a press release each time; the pressure at your tap is decided by the pipes.
It would be too neat to say the upgrades give the majority nothing, and the neatness would be a lie. The newer models are more reliable in ways that matter every day even when the output is flat: they decline to answer when they do not know, where their predecessors confabulated cheerfully; they hold longer documents in coherent focus; they require less of the quiet, exhausting human verification that has been this technology’s hidden cost. A model that hallucinates less is a real improvement in everyone’s Tuesday, even if the prose it produces is the same prose. But it is a gain in trust rather than in capability, and it accrues slowly, version on version. Less a leap than a settling.
Nor is any of this an argument against upgrading. The upgrade is usually a configuration change: one line, a few minutes, often a price cut. And refusing it on principle would be its own kind of theatre. The model and the infrastructure are complements, not rivals. Take the new model; it is close to free. The error is not in upgrading. The error is in mistaking the upgrade for progress, in believing that the organisation moved forward because a dropdown menu did.
The correction is not to stop watching the livestreams. It is to recalibrate what they are: a quarterly weather report for a climate most of us do not live in. For the engineers and the specialised few, each release genuinely changes what is possible, and their excitement is earned; everywhere else it depreciates by the following Thursday, and the quiet disappointment gets filed under expectations rather than under diagnosis. What changes Monday morning for the rest of an organisation ships to no announcement at all: a connector that gives the assistant access to the ticketing queue and alters the day of forty people in a way the launch never could, a skill that makes the first interaction useful instead of generic, a workflow rebuilt around what the tools already do. That infrastructure compounds, quietly, release after release. The people refreshing the stream are not foolish; they are watching something genuinely remarkable. It is just, for most of them, not addressed to them.