An argument is circulating through boardrooms, and it flatters you before it frightens you. You, the reader, rank among the best in the world at using artificial intelligence; your colleagues do not, and they never will. The chart that accompanies the argument looks like a barbell: a small weight of brilliant users at one end, a large dead weight of people who never log in at the other, and nothing in between. The distribution appears at every company, the argument goes, at fifty people or five thousand, and no rollout ever changes it. The prescription follows from the shape: stop training, and hide the machine inside the software people already use.
I run AI enablement for a company of about a thousand people, which means I hold the usage data this argument makes such confident predictions about. The predictions fail. More interesting than the failure is why, because the argument commits a specific economic error, one with a long history, and the error points to a better plan.
Test the shape first. I built the exact tier split the argument predicts on our own seat-by-seat usage data. The top matches: about one in ten of our people qualify as power users, running the tools across every surface at heavy volume. Below that, the prediction collapses. The argument requires roughly seventy per cent of seats to sit dark. Ours: twelve per cent. In place of the missing dead weight sits the tier the argument says cannot exist: a majority of the company, fifty-two per cent, doing real work in these tools on an ordinary working day. The distribution is a pyramid with a thick middle, the one shape the barbell thesis rules out.

One could object that my company is the outlier. Grant the objection and the argument still loses, because it claimed invariance. A distribution that holds "every time, at fifty or five thousand" admits no counterexamples. One pyramid ends the law, and demotes the barbell to what it always was, a description of the rollouts that went badly.
Our data carries a second, more damaging fact. The depth of use is not stable across the company. Our analysts use the tools deeply at nearly twice the rate of our finance team; one country in our footprint runs decades ahead of another, on identical software and identical licences. Economists have a name for this pattern. Within any industry, measured productivity varies enormously across firms doing the same thing with the same technology: in United States manufacturing, the plant at the ninetieth percentile produces roughly twice as much from the same inputs as the plant at the tenth. Two decades of work, much of it by Nicholas Bloom and John Van Reenen, traces a large share of that dispersion to management practice. A gap that swings with the function, the manager and the country is a management variable, and management variables move. The barbell asks you to read a management variable as a constant of nature, then charges you for the fatalism.
The deeper error sits in the unit of analysis. The argument prices each worker with a single number (an "AI skill level") and sorts the company by it. The experimental record of the past three years says the number does not exist. The return to AI is a joint property of the person and the task, and the task usually dominates.
Consider what happens at the bottom of the skill ladder. In the largest field study to date, economists tracked 5,179 customer support agents as an AI assistant rolled out across their contact centres. Productivity rose fourteen per cent on average. And the average concealed the finding. Novice agents gained thirty-four per cent, while the most skilled gained close to nothing, and on some quality measures lost ground. Agents with two months of tenure began performing like untreated agents with more than six. The tool had absorbed the working habits of the best performers and handed them to the weakest. A randomised experiment on professional writing tasks, published in Science, found the same signature: completion time fell forty per cent, quality rose eighteen per cent, and the gap between the strongest and weakest writers narrowed. A field experiment on 758 consultants at a major strategy firm completed the triptych: on tasks within the model’s competence, the bottom half of consultants improved forty-three per cent against seventeen for the top half.

Read those three results together and the barbell inverts. In bounded work, AI behaves as a substitute for skill, and the people the argument writes off capture the largest gains. Earlier waves of information technology complemented skill; the economists running the support-agent study note, pointedly, that their result cuts against that entire literature. The tool compresses the distribution the argument insists it stretches.
Now the top of the ladder. The consulting experiment carried a second condition, a task chosen to sit outside the model’s competence. There, consultants with AI access performed nineteen percentage points worse than colleagues working unaided, misled by fluent, confident, wrong output. A 2025 randomised trial by the research group METR pushed the point further. Sixteen veteran open-source developers, working in codebases they had known for an average of five years, took nineteen per cent longer on real tasks when allowed to use AI, and estimated, afterward, that the tools had made them twenty per cent faster. Set that against a controlled trial in which developers given an AI assistant finished a well-scoped, greenfield task fifty-six per cent faster, and the measured effect on a single profession runs from plus fifty-six to minus nineteen. The researchers behind the consulting study coined the right term for this: a jagged frontier. Tasks of similar apparent difficulty fall on opposite sides of the machine’s competence, and the boundary is invisible from the outside.
An economist reads this record and reaches for an old idea: the production function differs by task, so the return to skill differs by task. Where a task is bounded (a support ticket, a routine brief, a standard campaign email), the machine clears the ceiling for everyone, and further human skill earns no return past it; diminishing returns arrive fast. Where a task is open-ended (novel engineering, ambiguous analysis), there is no ceiling to clear, a harder problem always waits, and skill compounds. The frontier premium, the gap between a company’s best user and its median one, is therefore wide in engineering, moderate in analysis, and narrow to trivial across most operational functions. Collapse every role into a single distribution and you produce a statistic that describes nobody: a weighted average sold as a destiny.

The same logic disposes of a related claim: that each model release widens the chasm because the skill floor rises. For the bounded majority of corporate tasks, capability past the ceiling is unpriced surplus. A marketing team drafting campaign copy gains nothing from the model’s new talent for proving theorems. Most releases, meanwhile, lower the floor: chat replaced prompt engineering, agents replaced chat. The gap widens only in the unbounded domains, and the argument’s habit is to generalise from exactly those.
The seventy per cent who never open the tool deserve a second look, because the argument treats them as a wound AI inflicted. Organisational slack predates the transistor. Large firms have carried underutilised labour for as long as economists have measured them, and the bigger the firm, the smaller each worker’s marginal product and the easier the slack hides in the aggregate. A one-person company with an idle worker produces nothing; a hundred-thousand-person company with fifteen per cent idle produces almost exactly what it produced before. Note where the argument aims its alarm: at the giant enterprise, the one place slack is cheapest to carry. The dormant seat is not a new pathology. It is old slack wearing a new and unusually legible marker, and the marker mostly identifies the people who were already coasting before the software arrived.
The adoption statistics the argument mocks deserve the mockery, though the diagnosis is wrong. Eighty-eight per cent of organisations now report using AI somewhere; a single-digit share report any meaningful profit impact. The gap between those two figures is Goodhart’s law in operation: a login count became a target, so it stopped measuring anything. The honest response to a bad metric is a better metric (depth of use, share of workflows touched, throughput), and the honest reading of "we rolled it out and nothing got faster" comes from growth economics rather than from despair. Erik Brynjolfsson and his coauthors call it the productivity J-curve: general-purpose technologies depress measured productivity for years while firms build the intangible complements (the new processes, roles and workflows) and only then pay off. Electrification behaved identically. Paul David’s famous study found that American factories took some three decades to convert cheap electric motors into productivity, because the gains lived in rebuilding the factory floor around the technology rather than in the motor. A flat throughput line two years into a rollout is the oldest finding in the economics of technology. It indicts the complements, never the workforce.
One piece of the argument does describe my company: the best users’ methods fail to diffuse. We hold hundreds of internally authored skill files; the single most adopted one reaches thirty people. The argument explains this with a conspiracy. Power users hoard their edge, the story goes, because the gap is their leverage. My data cannot rule that out somewhere. It rules it out here, where the authors publish enthusiastically and the recipes still sit unadopted.
Economics supplies the duller, correct explanation. What a great user knows is tacit, the kind of knowledge David Autor, following Michael Polanyi, summarises as "we know more than we can tell." Codifying tacit knowledge is expensive, the benefit accrues to everyone except the codifier, and no one is paid to do it. Inside a firm this is a textbook underprovided public good: an incentive gap, not a character flaw, and incentive gaps have owners. A manager who sees one report moving at twice the pace of the team has one job that week: extract the method and make it the team’s method. The shared library the argument sells as its remedy is worth building, and I built one. It moves the artefact. It cannot move the judgment of when to reach for the artefact. And the argument itself concedes, a few paragraphs earlier, that this judgment, knowing which workflows should never touch a model and which should be automated end to end, is ninety per cent of the game. The remedy on sale addresses the ten.
The plan that survives the evidence runs two tracks in parallel, each sized to its own economics. Track one serves the unbounded work, where the frontier premium is real and wide. Train there, in engineering and analysis, and treat the training as diagnostic, because you cannot know who occupies your top slice until everyone has been given a serious chance to reach it. Give that top slice distribution, a place to publish, and make diffusion a line in every manager’s job, since tacit knowledge moves through managers or it does not move at all.
Track two serves the bounded work, and here the argument and I arrive at the same destination while disagreeing about what it proves. Most people are not in the market for a tool; they want the work done. For high-volume, repetitive workflows, build the automation into the systems people already use, so the work completes without asking thousands of employees to master a craft that pays them nothing in their own function. The strongest field evidence points exactly here. The support-agent study, the largest measured win in the literature, involved no prompt training at all; engineers embedded a tool in one workflow, and the least skilled workers gained the most. The gains came from the build. Companies with engineering organisations should staff this as an internal forward-deployed function: engineers who sit inside the business, find where the work actually happens, and wire agents into it. Companies without one should buy that capability in, and should weigh the fee against what an eight-figure licence commitment returns when it lands on unrebuilt workflows.
Report to your board accordingly. Drop the adoption percentage. Report the share of work performed manually, the share performed with assistance and the share fully automated, broken out by function, and watch the second and third numbers.
The barbell, in the end, is a mirror. It reflects a company’s own slack, its own unbuilt complements and its own managers back at it, and then invites the viewer to call the reflection physics. The experimental record, the productivity history and one mid-sized company’s usage logs return the same verdict. The gap is real, uneven, and governed by the task. Nothing about it is a law. Stop admiring the reflection, and build.