In short
The true cost of AI is becoming impossible to ignore, pushing organisations to make smarter decisions about which models they use, where workloads run, and when more expensive intelligence is actually worth paying for.
I read Jon Morra’s AdExchanger piece, “Marketers, The Free Token Lunch Is Over. Is Your Stack Ready?” this week, and I have kept coming back to it. He put a clever name to a conversation that now comes up almost every day in our business: the true economics of AI are finally showing up on the bill.
His argument is straightforward. The habit of sending every task to the biggest model available was never going to survive enterprise-scale usage. Marketing leaders should find that clarifying rather than alarming. It forces us to understand what we are buying and to make better choices about where these systems genuinely help.
My view comes from the intersection of technology, analytics, and agency operations. I see how the economics land in day-to-day marketing work, where quality, speed, risk, and margin all meet. From that vantage point, the organisations that do well will develop judgement at the task level. Spending targets by themselves tell us very little. Capability, quality, and cost need to fit the work.
The mixed message inside AI adoption
Most organisations are broadcasting two messages at once. The first: everyone should be using AI, every day, for everything. The second: please stop wasting tokens. Both messages are reasonable, but side by side they can be baffling. Teams need enough context to reconcile them.
A simple classification task or first-pass draft may run perfectly well on a lightweight model. Work that requires complex reasoning, deep domain knowledge, or a very low tolerance for error can justify more capable and more expensive systems. The hard part is knowing which situation you are in. Reasoning models can consume substantial internal tokens before producing a visible answer, and the usage can vary considerably across models and settings.
Once people understand that trade-off, they usually make better choices. They can decide when a stronger model is buying a material improvement and when it is simply adding cost. Treating consumption cost as a finance secret, or surfacing it only when an invoice surprises someone, leaves teams without the information they need. Transparency supports adoption and efficiency at the same time.
We’ve seen this film before
The transition from generous flat-fee access to more visible consumption pricing was foreseeable. Early pricing lowered the cost of experimentation, which helped the market move quickly. At scale, providers and enterprise buyers both need a closer relationship between usage and cost.
If you are of a certain vintage, the moment feels familiar. Many of us remember watching our mobile minutes and waiting for off-peak hours before making a call (if you know, you know). Those tariffs trained us to treat airtime as scarce. Competition, capacity, and better economics eventually made that anxiety disappear for most people.
AI compute may follow a similar arc. We are still in the watch-your-minutes phase, and budgets get blindsided when teams pretend otherwise. Unit costs may continue to fall, but usage often rises faster than the price of a token falls. More use cases, longer contexts, additional reasoning, retries, and agentic workflows can easily absorb the savings.
Morra’s most useful point is the idea of cost-of-pass: the expected cost of getting an acceptable answer. Per-token price is only one input. A useful calculation also includes retries, latency, human review, and the cost of a failure. That framing makes model routing an operating discipline with direct implications for quality, margin, and risk.
The price and value conversation is back
This is where the issue lands most directly for agencies, and for knowledge work more broadly. AI-assisted work can take fewer human hours, and clients should expect to share in genuine efficiency gains. At the same time, those hours can now contain deeper analysis, broader coverage, faster iteration, and work that was not feasible at a sensible price a few years ago. The hours are fewer; what fits inside them is bigger.
The cost base is changing as well. Behind an AI-assisted deliverable sit compute, tooling, orchestration, governance, evaluation, and the senior talent required to direct the work responsibly. Those costs will not rise evenly, and human expertise will remain the largest line item in many engagements. Still, “the model did it” is a poor proxy for both the cost of the work and the value of the result.
A healthier commercial conversation starts with outcomes, service levels, and risk. We should be willing to show where AI reduced effort and where it introduced new cost. Clients can then evaluate the resulting scope and outcome without using hours as the only shorthand for value. If everyone is looking at the same honest ledger, this stops being a fee negotiation and starts being a shared win.
I think a local compute renaissance is coming
Here’s the hot take I am most interested in. Local and private compute will regain a meaningful place in enterprise AI stacks. Open-weight models are improving quickly, and the software for running them on Apple silicon and other accessible hardware has matured. A small stack of Mac Minis (or their equivalent) in an office or data centre can now work around the clock, a kind of digital support crew that keeps data and operating control closer to the organisation.
The purchase price is only one line in the calculation. A credible break-even model also includes utilisation, power, support, security, model maintenance, reliability, and the cost of idle capacity. Stable internal workloads are the best candidates, especially when privacy matters and the required quality is well understood. Spiky demand, fast-changing requirements, or problems that need the strongest available models will continue to benefit from cloud services.
High, steady utilisation changes the calculation. For internal analysis, drafting, classification, and the daily connective tissue of knowledge work, owning some capacity will become competitive against renting every token. I expect many organisations to arrive at a hybrid stack: local or private models for repeatable work, with frontier APIs available when a task genuinely benefits from them.
That will make model choice a strategic capability. Teams will need to know how much intelligence a task requires, how they will measure a good result, and where the work should run. The answers will vary by workload, which is exactly why a single-model strategy will become harder to defend.
What mature adoption looks like
None of this dampens my enthusiasm for AI. It makes me want organisations to approach adoption with the same seriousness they bring to any other operating capability. Teams need visibility into usage, reliable evaluations, sensible routing rules, and permission to choose a smaller tool when it is the right tool.
The next phase will ask more of us than enthusiasm. We will need to understand the bill, design around it, and decide where expensive intelligence earns its keep. The free lunch was fun while it lasted. Knowing exactly what we are paying for, and what it is worth, feels like progress to me.
Contributing Experts
Mentioned in this article
Mentioned in this article
Proove Analytics
Unlock actionable insights through web analytics, data science, and business intelligence powered by our center of excellence, Proove Intelligence.
Learn More