Skip to content
Blog

Craft

AI automations in an agency: the ones that return nothing

An automation returns nothing until the time it saves has a written destination. Three tests before you install the first one.

The ones whose saved time has no written destination, and that is almost always half the list. An automation that works perfectly but whose returned hours get reabsorbed by the same tasks produces no measurable result, and it will nonetheless be counted as a success in the year-end review.

Lists of automations to put in place are useful and good ones exist, starting with the one Cobalt keeps for agencies and IT services firms, ranked by impact and ease of implementation. This article does not redo it: it deals with the question asked immediately afterwards, which is which of them produces anything, and on what condition.

Where does automation actually settle?

Where volume is high and judgement is low, and the 2026 split says so without ambiguity.

The Aptitude Research and iCIMS surveys published in April 2026 give 58% adoption on application screening, 54% on candidate communication, 50% on assessments and 46% on sourcing. The comment alongside those numbers is worth more than the numbers: adoption is heaviest where the flow is high and the process repetitive, lightest where evaluation requires judgement.

Size matters as much as function. The SHRM survey run in 2026 across nearly nineteen hundred organisations gives 60% adoption above five thousand employees and 33% in firms under a hundred. That is not a technology lag, it is a threshold: automation pays back on volume, and an agency of fifteen people has neither the same flow nor the same arithmetic. Copying a large account’s list of use cases is the surest way to automate what will return nothing at your scale.

The figure to read backwards

You will commonly read that teams using these tools seriously gain about a dozen hours per recruiter per week, with time-to-fill around fifteen days. The measurement comes from a vendor, across its own user base, and I give it with that reservation. But the reservation is not what matters here.

What matters is the condition attached to that figure, almost always in small print: it covers teams that redesigned their process around the tool. In other words it is not the automation producing the gain, it is the reorganisation the automation made possible and somebody decided on. Teams that installed the same tools on the same process as before do not appear in that figure, and they are the majority.

That reading also explains why so many projects stop. Gartner forecasts more than 40% of agentic AI projects cancelled before the end of 2027, and we have detailed elsewhere why these projects fail: almost never for a technical reason. A tool that works, laid over a process nobody wanted to touch, produces exactly what was asked of it and none of what was expected from it.

The three tests before automating anything

Three questions, in this order, and one negative answer is enough to push the task further down the queue.

Is the result faster to verify than to produce? That is the criterion that sets the order, and we have argued it for a long time: you delegate what is verifiable first, not what takes the most time. Interview preparation is checked in thirty seconds. A ranking of two hundred applications is not checked at all, and the trust placed in it rests on nothing.

Is the error reversible? A badly written summary is corrected, a message gone to a client is not. That is why we separate everywhere what stays inside from what goes out, and why nothing outbound leaves without a person releasing it, including a routine nudge.

Does the returned time have a named destination? This is the test nobody applies and the one that decides the outcome. Two hours returned to a business manager become, absent a decision, two more hours on the same tasks. Writing down what those hours must become, before launching the automation, turns a theoretical gain into an observable change: three more consultant check-ins a week, ten dormant accounts called, five candidates called back the same day.

Three automations that rarely return what they promise

They are not bad in themselves, and that is the problem: they work technically and disappoint economically.

Outbound messaging at volume. Multiplying candidate approaches or client nudges increases production without reducing anybody’s work, and moves the load to whoever receives it. That is the mechanism of work that looks finished: the gain is measured on the producer and declared, the load scatters across the others and never is. In a trade where silence already drives candidates away, generic volume worsens the reputation it was meant to repair.

The automatic score on a profile. A score reads in one second, does not say what it measured and therefore cannot be contested; it simply applies. We have written why it is the wrong instrument and what replaces it, namely a sentence that can be refuted.

The summary of everything, sent to everybody. Meeting notes, file synopses, weekly recaps: each looks free and each creates a verification read for five people. The rule that limits the damage is easy to state and tiresome to hold: a document produced automatically has a named recipient and an expected action, or it is not produced.

How many at once?

One, and the answer disappoints every time I give it.

The argument is not caution, it is measurement. Five automations launched in the same month produce a quarter in which something improved without anybody able to say which one is responsible, or which cost more than it returned. You get a global result and no usable information, which leaves you either keeping everything or stopping everything, and that is how most abandonment decisions get made.

One alone, launched with its time destination written in advance and an indicator chosen before the start, produces a clear answer within six weeks: it worked, it did not work, or it worked for two people out of twelve and you need to understand why. That pace looks slow on an annual plan and is faster in practice, because it never starts again from zero.

It has another merit, less avowable and quite real. One automation at a time means one conversation at a time with the teams who will use it, and adoption is the only factor that really decides the outcome. An excellent tool nobody opens is worth exactly zero, and that shows on no dashboard until the quarter ends.

So where do you start?

With what is memory rather than judgement, because it is the one part of this trade that delegates without losing anything.

Knowing who has been waiting for an answer and for how many days, which consultant is entering their sixth month on assignment, which client account has invoiced nothing in eighteen months, preparing a meeting with what was said last time: those tasks pass the three tests effortlessly. The result is checked at a glance, the error is reversible because nothing goes out, and the destination of the returned time is obvious since it is the work itself.

That is where we started ourselves, and not out of virtue: the first things Balt could do were memory tasks rather than writing tasks, because they are the only ones a user can tell are wrong in three seconds. A list of forgotten follow-ups is checked at a glance; a well-written message is not checked at all.

It is also, incidentally, the part automation lists cover least, because it is spectacular in no demonstration. A tool that tells you “you have not called these eleven people back” does not look like artificial intelligence. It changes a month of work more than any message generator.

That leaves the question that follows immediately, and decides half the failures: once the task is automated, who is accountable when it runs alone at three in the morning and fails in silence. Nobody watches that run, and that is exactly why its owner has to be decided now.

Frequently asked questions

Which automation should an agency start with?

The one whose result can be checked at a glance, not the one that takes the most time. Interview preparation from a file is verified in thirty seconds; automatic ranking of two hundred applications cannot be verified at all, and trust is built on the first case.

How much time does an automation actually save?

Reported figures go up to twelve hours per recruiter per week, but they cover teams that redesigned their process around the tool. Without that redefinition, the time returned is reabsorbed by the same tasks and appears in no indicator.

Why do large companies automate more?

60% above five thousand employees against 33% in firms under a hundred, because automation pays back on volume. An agency of fifteen people does not have the same arithmetic, and copying a large account’s use cases is the surest way to automate what returns nothing.

Which automations should be avoided?

Those that increase outbound message volume without reducing anybody’s work. They move the load to whoever receives, candidates and colleagues alike, and that transfer gets counted as a gain because it is measured only on the side where it shows.

Sources

  1. Pin, AI Adoption in Recruiting: The Complete 2026 Industry Report (Aptitude Research, iCIMS and SHRM data)pin.com
  2. Cobalt, 20 automatisations IA à mettre en place dans votre cabinet ou ESNcobalt-ia.com
  3. Gartner, Over 40% of Agentic AI Projects Will Be Canceled by End of 2027gartner.com

Read next

We pay you to work less.Get your €100 now.

Join the waitlist.

Leave your email address and we will let you know as soon as Balt can join your team.

Already 247 staffing firms on the waitlist