— Notes from Mallin

Nobody has a control group

Ask almost any executive whether AI is working in their company and you'll get a confident answer. Ask what it's being compared against and the conversation gets vague fast.

That vagueness is the whole story.

The one honest experiment

In 2025, METR ran the study almost nobody else bothered to run: an actual randomized controlled trial. Sixteen experienced open source developers, 246 real tasks from codebases they already knew well, each task randomly assigned to either allow AI or forbid it.

The developers predicted AI would make them 24% faster. It made them 19% slower.

That's not the interesting part. Here's the interesting part: after finishing the tasks, after living through the slowdown, they still estimated that AI had made them about 20% faster. The experience of being slower did not register as being slower. It registered as speed.

The authors are careful and say plainly that this doesn't prove AI fails to speed up most developers. Sixteen people, one domain, mature codebases, limited time on the tools. Take their caution seriously.

But the perception gap is not a footnote about software engineering. These people were measuring themselves in the moment, at the thing they do professionally, and were wrong by nearly forty points in the wrong direction. They were not fools and they were not lying. They had exactly the evidence your organization is running on right now, and that evidence pointed the wrong way.

The number that should have ended the argument

Meanwhile, the aggregate keeps refusing to show up. PwC's 29th Global CEO Survey found that more than half of chief executives, 56% of them, realized neither revenue nor cost benefits from AI in the past twelve months. One in eight got both.

The reflex reading is that AI is overhyped. I don't think that's it, and I don't think the opposite reading holds either.

Look at what those CEOs were actually asked to do: report a financial effect they mostly have no instrumentation to detect. Almost nobody set a baseline before deployment. Almost nobody held anything back for comparison. Half the value shows up as work that didn't need doing, which is invisible by construction. So 56% of CEOs saying "no measurable benefit" and their own teams saying "this is transformative" are perfectly compatible statements. Both are guesses.

We have run the largest reallocation of enterprise software spend in a decade, and the primary instrument has been how fast the work feels.

You are measuring the thing the product is best at faking

Here is why the feeling can't be trusted, mechanically.

Perceived productivity is built from friction. Effort feels like cost, and its absence feels like progress. These tools are extraordinarily good at removing the sensation of effort. Instant output, fluent prose, no blank page, no waiting. That is the single most optimized property of the entire category.

So when you evaluate on how it felt, you are measuring the one dimension the technology is guaranteed to win, whether or not anything downstream improved. The blank page got filled in four seconds. Whether the thing on the page was right, whether it took three more meetings to unwind, whether the customer noticed. None of that is in the feeling.

Selling has an acute version of this, because its favorite metrics were always activity metrics. Emails sent, accounts touched, calls logged, pipeline generated. Those numbers were never good, but they were once weakly informative, because producing them cost human hours. That cost was the entire signal. Remove it and pipeline generated stops measuring a team and starts measuring a subscription.

If your outbound volume tripled and your closed won number didn't move, you did not get more productive. You got more expensive, with prettier dashboards.

Buy a control group

The prescription is unglamorous and cheap, and almost nobody does it.

Hold something back. Pick one workflow. Give the tool to half the team for six weeks and not the other half. You will get an argument about fairness. Run it anyway, and rotate. Six weeks of one honest comparison is worth more than two years of testimonials, and you will be the only company in your market that actually knows.

Write the baseline down before you deploy, not after. Cycle time, win rate, whatever you claim will move. A number recalled after the fact is not a measurement; it's a story with a number in it.

Measure outcomes, never volume. When production is free, every production metric decays into a measure of your vendor's throughput.

Two predictions I'll stand behind. Within two years, AI return on investment becomes a formal audit category, and a meaningful share of the gains currently reported in board decks will not survive contact with a real baseline. And the companies that come out ahead won't be the ones that adopted earliest or spent most. They'll be the boring ones that ran the comparison, found that three of their eight deployments did nothing, killed those three, and put everything behind the two that measurably worked.

The competitive edge in this cycle isn't access to the models. Everyone has that.

It's being one of the few organizations that can tell the difference between working and feeling like working.