Craft
Why 40% of AI agent projects will be abandoned
Gartner expects over 40% to be abandoned before the end of 2027, for three reasons: no measurable ROI, no governance, no integration. All three are preventable.
Gartner expects more than 40% of agentic AI projects to be abandoned before the end of 2027, for three named reasons: no clear ROI, no governance, no integration.
The word to read in that figure is “abandoned” rather than “failed”, and the distinction carries the whole subject. These projects work technically. They are stopped because nobody can say what they return, nobody knows who answers when they get something wrong, and they do not reach the tools where the work happens.
All three are preventable, and none requires technical skill.
First cause: an ROI nobody can read
The trap is measuring the wrong thing, and “time saved” is the metric everyone picks and nobody can ever establish. Saved against what, exactly? A business manager does not time their follow-ups, so there is no baseline to compare against. Six months later the committee asks for the number, nobody has it, and the project becomes an expense with no justification.
What can be measured is what did not exist before. The number of follow-ups that actually went out this month, when a third of them never used to. The number of roll-offs spotted at sixty days instead of thirty. The number of interview write-ups actually written rather than promised.
Those numbers have two virtues: they can be counted without interviewing anybody, and they convert into euros without a heroic assumption. A roll-off anticipated by three weeks, on a consultant costing around €5,480 a month, is a figure nobody argues with in a committee.
Fix that metric before you start, not at review time. A number chosen afterwards is always suspect, and rightly so.
Second cause: nobody knows who answers
An agent writing to candidates raises a question the organisation has never had to settle: who is responsible for what it sends?
While it goes unanswered, there are two outcomes, equally fatal. Either everything waits for approval, including what did not need it, and the cost of re-reading cancels the gain: the team concludes it “does not save time”, which is true in that configuration. Or nothing waits for approval, and the first clumsy message to a senior candidate ends the project in a single meeting.
The line that holds is mechanical: what leaves the company waits for approval, what stays inside goes on its own, what closes a door is decided by a person. That is the subject of a whole article, because it is the decision that determines whether the project survives its first incident.
And you need a log. With no record of who approved what, an incident ends in “the AI did something odd”, which teaches nobody anything and costs the whole team its confidence over what may well have been a badly phrased brief.
Third cause: the agent does not reach the work
This is the most ordinary failure and the most expensive one. An agent living in its own interface gets delegated almost nothing, because you have to think of it, go there and re-type the context every time. By the third time you do the task by hand: it is faster.
The corollary is less obvious: an agent well integrated into a disorderly system of record is worse than useless. It produces sound reasoning over wrong data, with a machine’s confidence. An impeccably argued shortlist drawn from a pool whose availability dates from last year wastes more time than it saves, precisely because it is credible.
Hence the counter-intuitive order of operations: check the state of your data before evaluating an agent. If your roll-off dates are not current in the ATS, no agent will anticipate them. It is also why the real question about integrations is not how many there are, but whether the four or five tools you open every morning are among them.
The fourth cause, which Gartner does not name
Impatience.
Break-even on an industrialised deployment lands between 4 and 9 months. Adoption takes three to four weeks just to start, and the habit of delegating builds more slowly still.
A committee judging at three months therefore stops before it has the answer. And it stops for apparently good reasons: the numbers are thin, usage is patchy, somebody has an anecdote about an error. All of that is normal at three months and says nothing about month nine.
Set the evaluation horizon at kick-off, in writing and alongside the metric, because it is above all a protection against yourself.
The pilot trap
There is one near-guaranteed way to fail, and it is running an isolated pilot, on reasoning that sounds perfectly prudent. Take a narrow scope, with nothing at stake, two volunteers, to “see how it goes without risk”. Six weeks later the demo has gone well and nobody knows what to do with it, because a pilot with nothing at stake produces no defensible number, and use without pressure proves nothing about use under pressure.
A small scope is a good idea; small stakes are not. Take a narrow but real process: one business manager’s follow-ups on their own accounts, not “follow-up testing”. In six weeks you get a number worth arguing about in a committee, and a team that genuinely needed it to work.
The other variant of the trap is giving the pilot to the most curious people. They will succeed, and they will prove nothing: they are the ones who would have found a way to save time regardless. The interesting question is what happens with someone who is not interested.
What to write down before the first connection
Three things, on one page, before opening an account.
The metric, and its value today. “How many follow-ups went out last month” has a measurable answer right now; without that starting point, the same number in six months means nothing.
The approval line, in one sentence, and who applies it. Not a governance document: a sentence the team remembers.
The horizon, and what would count as failure. A date, a threshold. Deciding it cold protects you from the March meeting where somebody tells an anecdote about an error.
One page. The time spent writing it is the best investment in the project, and it is precisely what four projects in ten did not do.
The only adoption metric worth tracking
Not logins, not request volume. The number of requests nobody suggested.
Usage statistics on a tool management asked people to use measure compliance. An unprompted request measures adoption, somebody had a problem and thought of the agent before thinking of doing it by hand.
The most reliable sign is simpler still: the day someone on the team is annoyed that something is not possible yet, since people only get annoyed about things they were counting on.
What a project that worked looks like at month nine
Worth describing, because the picture is less dramatic than people expect.
Nobody talks about the agent any more. It is not on the agenda, there is no dashboard anyone opens, and the person who championed it has moved on to something else. The follow-ups go out. The write-ups exist. Somebody noticed a roll-off in March that would have been noticed in May.
The team cannot tell you how much time it saves, and that is fine: the number you fixed at kick-off can, and it moved. What they can tell you is which two things they would never go back to doing by hand.
The failure mode looks different in one specific way: at month nine a failing project still has meetings about it. Attention is the symptom. A tool people are still discussing is a tool that has not become part of the work.
What that changes about starting
Three decisions, taken before the first connection, and written down.
Which number you will look at in six months, and how you count it today to have a baseline. Who approves what, on the “leaves / stays” line. Which tools the agent is plugged into, and whether the data in them is current.
None of the three is technical. Which is exactly why they get skipped, and why four projects in ten stop while still working.
Frequently asked questions
Are these technical failures?
Almost never. The three causes Gartner names, unreadable ROI, absent governance, missing integration, are organisational. The project works in a demo, ships, and quietly dies for want of use. That is an abandonment, not a breakdown.
How long before judging an agent project?
Six months at minimum. Break-even on an industrialised deployment lands between 4 and 9 months, and adoption takes three to four weeks just to start. A committee deciding at three months stops before the answer, every time.
Start small or go big?
Small, but on a real process rather than an isolated pilot. A narrow, real scope gives you a number worth arguing about in six weeks; a laboratory pilot gives you a demo nobody can turn into a decision.
Which metric should we track?
Unprompted requests: the ones nobody suggested. Usage statistics on a tool the team was told to use measure compliance, not adoption.
Sources
Read next
Product
AI agent or automation: which one for which taskAn automation follows rules, an agent makes decisions. The rule for choosing: if you can write the script, automate. If you can only describe the outcome, delegate.Delegation
Delegating to an AI agent: where to startStart with work you can check in ten seconds: follow-ups, interview notes, record updates. Here is the order that works, week by week.
