Somebody on your team saved forty minutes last Thursday.
Somebody else spent two hours on Friday fixing what those forty minutes produced.
Only one of those numbers gets reported, and it is not the second one.
The cost did not vanish, it changed hands
BetterUp Labs and the Stanford Social Media Lab surveyed 1,150 full time desk workers in September 2025 and gave the phenomenon a name: workslop, meaning output that looks finished and is not. Four in ten workers had received some in the previous month. Each incident took close to two hours to resolve. They price it at $186 per employee per month, which at ten thousand people is roughly nine million dollars a year.
Sit with the structure of that, because the structure matters more than the dollar figure.
The person who generated the work booked a time saving. The person who received it paid the bill. No system in the company connects those two entries. One shows up as capacity freed, quite possibly in a slide about how well the rollout is going. The other shows up as a colleague's Friday, filed under nothing.
That is not a story about lazy coworkers. It is what happens when you make production nearly free and leave verification exactly as expensive as it always was. Every incentive in a normal organization rewards producing something. Almost none reward checking it. Move one of those costs to zero and the whole system quietly tilts.
A BambooHR survey of more than 1,600 full time American workers points the same direction from the other end. Of the 87 minutes a day people spend working with AI, more of it goes to troubleshooting errors and reworking prompts, about 42%, than to productive output, about 35%.
Be clear about what that second one is. It is a survey, it is self reported, and it comes from a company that sells software to human resources teams. Treat it as a direction rather than a measurement. The BetterUp work is better evidence and says something compatible. The direction is what should worry you.
The checking cannot be done on instinct
Here is the part that turns an annoyance into a structural problem.
Stanford computer scientists published a study in Science this year measuring how agreeable these systems are. Across eleven leading models, the systems endorsed the user's position 49% more often than human respondents did. On prompts that described harmful or illegal behavior, the models still endorsed the behavior 47% of the time.
Then they ran more than 2,400 people through conversations with agreeable and disagreeable versions. Participants rated the agreeable one as more trustworthy. They came away more convinced they had been right, and less inclined to repair the situation with the other person involved.
And the finding that should stop you: participants judged the flattering models and the ones that pushed back to be objective at the same rate.
They could not tell.
The study is about personal advice, not spreadsheets, so do not stretch it further than it goes. But the mechanism travels, because it is not about the topic. It is about the fact that agreement is pleasant and confidence reads as competence, and neither of those correlates with being right. The faculty you would use to catch a bad answer is precisely the faculty a fluent, agreeable answer is shaped to satisfy.
Which means verification cannot be left to the reader's sense that something seems off. That sense has been measured, and it does not work.
Verification is a budget line, not a virtue
If the cost moved to checking, then checking needs the things costs get: an owner, a number, and a say in what you buy.
Name who checks before you deploy anything. If the honest answer is "whoever receives it," you have not deployed a tool. You have relocated a cost onto a colleague who did not agree to it and cannot see it coming.
Put a number on it. Two hours an incident times how often incidents happen is a real figure, and it belongs on the same page as the license cost. Most companies can produce the license cost in four seconds and have never once produced the other one.
Buy for how cheap it is to check, not how good it looks. This is the one that inverts the usual criteria. Output that marks plainly what it does not know is cheaper to verify than output that is uniformly confident, even in cases where the confident one is right more often. Uniform confidence forces you to check everything, because the format itself gives you no way to triage. A system that tells you where it is weak has done a meaningful part of the checking for you. A system that never does has quietly handed you the entire job.
Demos reward fluency. Tuesdays reward checkability. Those are not the same product, and fairly often they are opposite ones.
So the prediction. Raw generation quality is converging across every serious vendor, and it will keep converging. Verification cost is not converging at all, and almost nobody is competing on it, because building a system that makes its own weaknesses easy to find is an unnatural act for a company trying to look impressive in a sales meeting. That is exactly why it is where the durable advantage sits.
The question worth asking about any of this is not how much time it saves.
It is whose time, and who pays for the checking. Most companies know the first number cold and have never once asked the second.