Skip to content
Blog

Vision

Intelligence is becoming free. Where does the value go?

The cost of the same level of answer fell 280-fold in eighteen months. When intelligence costs nothing, what becomes scarce is context and the right to act.

To what cannot be copied: proprietary context and the right to act. That is the conclusion, and the rest of this article explains why it is not a slogan.

The numbers first, because they are under-appreciated and dramatic. Querying a model at GPT-3.5 quality cost $20 per million tokens in November 2022. It cost $0.07 by October 2024: a 280-fold drop in eighteen months. At GPT-4 level, from $30 to under $0.50. The observed trend, at constant capability, is roughly a fourfold reduction per year.

No industrial input has ever behaved like this.

What it destroys

The direct consequence is that the model stops being an asset. Choosing your model provider has become the least structuring decision in an enterprise AI project, which is a complete reversal from 2023.

Three things disappear with it.

The advantage of access. Everyone has the same capabilities within a few months, and the gap between the best model and the third closes faster than a sales cycle.

Justifying price by technology. “We use the best models” is no longer an argument, it is a description of everybody.

Demos. When anyone can produce an impressive demonstration in an evening, the demonstration stops discriminating, which explains part of the project abandonment rate and why the real causes of failure lie elsewhere.

What it makes scarce

A resource becoming abundant moves scarcity elsewhere; it does not remove it. Scarcity now sits in three places, and none of them can be downloaded.

Proprietary context. A model that cannot see your talent pool, your client exchanges and your assignment history is guessing. An estimated 85% of enterprise agent errors come from missing context, not from model shortcomings. The consequence is direct: your poorly maintained ATS cost you little while nobody read it, and becomes the limiting factor the day an agent uses it.

The right to act. Reading has become nearly free; acting has not. What is difficult, expensive and slow is getting a company to grant a system the right to write into its ATS, send a message in its name, touch its client relationship. That permission is not bought by the token: it is earned, it is revocable, and it is lost all at once.

Distribution. Being where the work happens: the inbox, the staffing tool, is worth more than having a better answer in a window nobody opens.

The optical illusion in the bill

A point almost everyone misses, with immediate budget consequences: unit price collapses while the bill goes up.

The reason is that the saving is immediately reinvested. An agent costing a hundred times less per token does not do the same thing a hundred times cheaper: it does something else. It re-reads, it verifies, it tries a second approach when the first returns nothing, it maintains a memory, it reasons longer before answering. Each of those is a genuine improvement, and each multiplies consumption.

Put differently, the falling cost of intelligence does not fund a saving, it funds quality. That is excellent news for the user and bad news for anyone expecting their bill to fall tenfold, and it is why what actually moves the cost of an agent is almost never the token rate.

What it changes for a buyer

Three concrete shifts in a scorecard.

Stop comparing models. “Which model do you use” has lost nearly all discriminating power. The question that kept it is: “what do you see of my data, and what are you allowed to do with it?”

Look at integration cost, not usage cost. The first is measured in months and does not fall; the second is measured in cents and falls by itself. Negotiating the second while ignoring the first is the most common purchasing error of 2026.

Assess reversibility. If the model is interchangeable, a vendor locking you in is not locking you on intelligence: they are locking you on your own connectors and configuration. That is the only angle from which a standard like MCP genuinely matters, and it is an exit criterion, not a capability one.

And a staffing firm, which sells intelligence by the day?

This is the uncomfortable question, and it would be cowardly to avoid it in an article published by a tool built for staffing firms.

A services firm sells human intelligence per unit of time. If artificial intelligence sees its price divided by four every year, what is the trajectory of a business model built on a day rate?

Our reading is that the pressure does not fall where people expect. It does not fall on the senior consultant who arbitrates, holds a client relationship and puts their signature on things: that person does exactly what a machine does not. It falls on the pyramid, meaning the model that consisted of billing several juniors to produce the volume the senior supervised.

The signal is already visible in consulting: firms that sold advice alone are losing ground to those that deliver, run and maintain what they designed. The deliverable stops being a document about the solution and becomes the solution. It is the same shift that makes per-seat pricing untenable for a vendor: you no longer pay for access to a capability, you pay for a maintained outcome.

Three consequences for a staffing firm, none of them theoretical.

Volume stops being sellable. What was billed because it took time will not be billed, because it no longer takes time. Nobody will buy three days of desk research.

What sells is the commitment. The responsibility to hold a scope, to answer for an incident, to still be there in eighteen months. A machine cannot carry it, for exactly the same reason a company cannot be 100% automated.

Training juniors becomes a problem you have to solve deliberately. The volume work was also the apprenticeship. If it disappears without being replaced by something else, the pyramid does not narrow: it empties from the bottom, and in five years you discover you have no seniors left.

Our position, and what it costs us

We build Balt on the assumption that intelligence will be free and trust never will be.

Concretely, that means investing little in what is about to commoditise, raw reasoning, and a lot in what will not: the quality of the context the agent receives, the precision of the permissions it is granted, the trace of what it did, and the line beyond which it asks.

That position has a commercial cost and it should be said. It produces less impressive demonstrations than an agent that was given everything, because half the work is invisible: what did not happen cannot be shown.

It also has a consequence for how a tool like ours should be priced. If intelligence tends to zero, charging for access to intelligence becomes untenable, and charging for work done rather than per seat stops being a preference and becomes the only tenable position five years out.

The prediction we will stand behind is this: within two years, nobody will sell AI. They will sell access, a scope and an accountability, and intelligence will be the cheapest ingredient in the package.

Frequently asked questions

How much has the cost of AI actually fallen?

A query at GPT-3.5 quality went from around $20 per million tokens in November 2022 to $0.07 by October 2024, a 280-fold drop. At GPT-4 level, from $30 to under $0.50. At constant capability the observed trend is roughly a fourfold reduction per year.

If models are equivalent, what differentiates a vendor?

Three things the model does not supply: access to the customer’s proprietary context, management of permissions and limits, and orchestration, knowing which model to call, when, with what in memory. Choosing the base model has become the least structuring decision in an enterprise AI project.

Should we wait for prices to fall further before starting?

No, because token prices are not what is holding you back. The dominant cost of a deployment is integration, getting the data in order and the time teams take to adopt it. Waiting for a discount on the cheapest line item delays the work that genuinely takes months.

Will falling costs make tools cheaper?

Not mechanically. Token cost collapses while consumption per task rises: an agent that reasons, retries and re-reads uses ten to a hundred times what a single query used. The bill for serious use falls far more slowly than the unit price.

Sources

  1. Shelly Palmer, Price per intelligence unitshellypalmer.com
  2. The GTM Newsletter, The token price collapse (and why AI costs still increase)thegtmnewsletter.substack.com
  3. The AI Corner, Marc Andreessen on why the AI moat is not the modelthe-ai-corner.com

Read next

€100 in credits when you sign up

Join the waitlist.

Leave your email address and we will let you know as soon as Balt can join your team.

Already 247 staffing firms on the waitlist