Skip to content
Blog

Delegation

The work that looks finished and is not

40 % of workers get it from a colleague, and each instance costs almost two hours to whoever picks it up. The cost does not vanish, it changes person.

Because the cost does not vanish, it changes person. The sender saves twenty minutes and says so; the receiver spends two hours working out what stands up in the document, and tells nobody. It is a transfer, not a gain, and it is invisible by construction.

The researchers at BetterUp Labs and Stanford’s lab who named the phenomenon measured it across 1,150 desk workers. 40 % had received it from a colleague, each instance cost almost two hours of rework, which comes to about 186 dollars a month per affected worker and close to 9 million a year in a 10,000-person company.

What the study actually measures, and what it does not

It does not measure model quality, and that is what makes it interesting. The text received is formally correct, articulate, error-free, often better presented than what the colleague used to produce. What is missing sits elsewhere: the decision is not taken, the client’s constraint is not folded in, and the three options are laid out with equal conviction.

The most reliable signal, when you are trying to recognise it, is the absence of a judgement call. A document that lays out without choosing hands the reader exactly the work its author was supposed to do, and hands it over wrapped in formatting that suggests the opposite. That gap between the appearance of completion and the substance is what costs the two hours.

Nor does it measure intent, and that has to be said because the moral reading of the subject is wrong. People who send this kind of work are not trying to cheat: they have too heavy a load, a tool that produces something presentable in thirty seconds, and no obvious reason to think the result falls short.

Why it damages trust more than the schedule

The most durable effect is not the delay, it is the judgement. 42 % of the people who received this kind of work judged the sender less trustworthy, and nearly half found them less reliable, less capable or less creative than before.

That degradation is asymmetric and irreversible in the short term. It takes one document to lose the reputation and six months of serious work to earn it back, which makes sending an unread text one of the worst ratios of immediate gain to deferred cost an employee can achieve.

For a staffing firm, the mechanism transfers straight to the client. A summary note that looks polished and does not survive reading says more than that its author moved fast; it plants a doubt about everything your team sends afterwards, including what was written line by line. It is the same reputational cost as the one a ghosted candidate circulates, pointed at the client this time.

The three places it shows up in a staffing firm

The consultant profile. It is the most exposed document, because it is produced in series, under pressure, and already reads like formatted text before any AI touches it. A profile generated from a CV without anybody having spoken to the consultant describes a plausible person, and the client finds out at their first technical question. The cost is not the document, it is the credibility of the next one you send.

The interview write-up. The trap here is subtler, because the text is often accurate. What is missing is the judgement: what the recruiter sensed, the hesitation on one question, the reason they recommend despite a weak point. A write-up without judgement forces the next person to redo the interview, and nobody ever counts that second interview as a cost of the first.

The tender response. This is where the money at stake makes the trade-off obvious, and yet it is the most tempting document to generate, because it is long and the deadline is short. A response covering twenty requirements at the same level of generality loses to one covering twelve precisely, and the time difference between the two is smaller than people assume.

What the three have in common is instructive: they are exactly the documents an AI writes best in form and worst in substance, because their value rests on information the model does not have. It is also why they delegate well as preparation and never as final output.

What this says about an agent, by contrast

The obvious objection is that we sell an agent that produces text, therefore a machine for manufacturing what this article criticises. It deserves a precise answer, because the difference is not in the model, it is entirely in the frame around it.

Work returned by a well-built agent can be checked at a glance, because it covers something checkable: a follow-up whose three lines you reread, a write-up you compare with what you have just lived through, a record update whose right value you know. That is why you delegate what checks quickly first, and not what takes the longest, even though the reverse looks more profitable on paper.

An agent also has to say what it could not do, and that is where the comparison is cruellest. The rushed colleague hands over an apparently complete document because nothing obliges them to flag their gaps; a properly built agent flags the missing information instead of filling it with a plausible phrase, and that behavioural difference is worth more than ten points of writing quality.

Finally, nothing leaves without a person releasing it. The rule looks heavy until you look at these figures: it is exactly the check that is missing when a generated text goes straight to a colleague or a client.

The rule that avoids it, and why it is unpopular

It fits in one sentence: whoever sends it answers for it, whatever tool wrote it. That rule existed before AI, nobody had ever needed to write it down, and it now has to be written because the tool made it possible to send without having read.

It is unpopular because it cancels part of the claimed gain. If you have to review seriously what the agent produced, you no longer save twenty minutes out of thirty, you save ten, and the dashboard looks worse. The dashboard was wrong, and the ten minutes are real.

There is a managerial variant that works better than a ban, and it consists of asking somebody to explain their document aloud in two minutes. An author who has worked the subject does it effortlessly; an author who pasted a model’s output notices it themselves by the third sentence, and there is no need to make a reproach of it.

This phenomenon explains part of the failure rate of enterprise AI deployments, and it is an angle you rarely read. Gartner expects over 40 % of agentic AI projects to be abandoned before the end of 2027, usually explained by the absence of measurable ROI, governance and integration. The first of the three is directly implicated here.

A ROI that counts only the time saved by the producer, without counting the time lost by reviewers, measures a transfer and calls it a gain. The project looks successful for two quarters, then somebody asks why the delivery times have not moved, and nobody has the answer because nobody measured the right side.

The practical lesson is short. Before deploying anything, decide who reviews, count what the review costs, and compare it with what the writing cost. If the arithmetic stops working once review is counted, that task was not the right one to delegate, and others are.

Frequently asked questions

How do you recognise this kind of work?

By text longer than it needs to be, impeccably structured, where every claim is plausible and none is verifiable. The most reliable signal is the absence of a decision: the document lays out options, does not choose, and leaves the reader the work its author was supposed to do.

Should AI be banned to avoid it?

No, and the companies that tried mostly produced clandestine use. What works is holding the author responsible for the content they send, whatever tool produced it, which was already the rule before AI and never needed writing down.

How is an agent different from a colleague pasting a generated answer?

A well-built agent returns work you can check in ten seconds, says what it could not do, and passes through a validation before anything leaves. A rushed colleague does the opposite of all three. The difference is not in the model, it is in the frame around it.

Does this cost show up anywhere?

Nowhere, and that is the heart of the problem. Time saved is measured by the producer and readily declared; time lost is scattered across several reviewers and never declared. A productivity assessment counting only the first shows a gain that does not exist.

Sources

  1. Harvard Business Review, AI-generated workslop is destroying productivity, September 2025hbr.org
  2. BetterUp Labs and Stanford Social Media Lab, Workslop: the hidden cost of AI-generated busyworkbetterup.com
  3. Harvard Business Review, Why people create AI workslop and how to stop it, January 2026hbr.org

Read next

€100 in credits when you sign up

Join the waitlist.

Leave your email address and we will let you know as soon as Balt can join your team.

Already 247 staffing firms on the waitlist