Blog

Why AI cost optimisation expires

Connect, optimise, done assumes the world holds still afterwards. It moves on three independent fronts, and only one of them is widely watched.

7 August 2026 · 8 min read

The standard AI cost engagement has a clean shape: connect to the providers, analyse a few months of spend, recommend a set of changes, leave. It is an honest shape, and for a market that holds still it is the right one. Nobody pays a consultant to keep re-checking a number that cannot move.

So the only question that matters is whether this market holds still. It does not, and it fails to in three separate ways that have nothing to do with each other. Any one of them can raise your bill while the other two do nothing at all.

Front one: the market re-prices itself every few weeks

Model pricing does not drift, it responds. A cheaper competitor arrives and the incumbent cuts. A flagship is quietly replaced by a stronger model at the same list price, which is a price cut expressed as capability rather than dollars. A provider that was the cheapest host for a given set of weights loses that position to another host running the identical model.

This is the front we can prove without qualification, because we keep an append-only ledger of every price move we have ever observed, with its date and its magnitude. Nothing in it can be edited after the fact. It is also the front with the sharpest implication for a one-time audit: the recommendations were not wrong. They were right on the day they were written, and the ranking they were derived from is regenerated by the market on a cadence the audit has no way to follow.

The practical version of this is unglamorous. A switch that saved twenty percent in March can be the more expensive option by June, not because anyone made a mistake, but because the alternative moved. Reversing a decision that has stopped paying requires knowing that it stopped, and that requires someone still looking.

Front two: your own stack grows into a bigger bill

The second front comes from inside the building. As a team's use of AI matures, the work it asks of a model gets heavier. Prompts accumulate context. Retrieval starts pulling more documents. A single-shot call becomes an agent that makes four tool calls before it answers. None of this is a spending decision in the sense that anyone approves it, and none of it involves the provider changing anything, but spend per task rises anyway.

We are going to be explicit about the evidence here, because it is the weakest of the three. We have not measured this across our own customer base. Our production history is not yet long enough to separate a workload that got harder from a workload that simply got busier, and a chart that cannot tell those apart is not evidence of anything. Treat this front as reasonable and widely reported, not as something we have demonstrated.

The market half of this argument is published, dated, and free to read

Open Intelligence

Front three: the setup you never touched changes underneath you

The third front is the one almost nobody monitors, because from the outside nothing happened. Same model identifier. Same prompt. Same code path. Same output, near enough that no test fails. And a different bill.

There are at least three mechanisms behind that. A provider can update the weights behind a stable model id. A reasoning model's default thinking effort can shift, so it thinks harder about the same question than it did last quarter. The scaffolding between your request and the model, the system layer you do not control and cannot see, can change how a request is assembled.

The reason this is invisible rather than merely unnoticed is worth understanding. On a reasoning model, a large share of what you are billed for is text the model generated while working out its answer and then discarded. It is reported in a different field from the answer, and it never appears in the response you receive. We caught one live: a model asked for a single word returned a single word, reported one output token in the field most cost tooling reads, and billed sixty-eight. Sixty-seven of those were thinking the model threw away.

What that proves is precise, and it is worth not overclaiming. It proves the bill for an unchanged request is free to move without the answer moving. It does not prove that it did move, for a particular workload, over a particular period. That is a different claim, and it requires the same task measured months apart.

We did not have that measurement, so we started taking it

From this month, eight fixed tasks run against six pinned models on the first of every month, and the token counts each provider reports are recorded in an append-only log. The tasks span the shapes real production traffic actually takes: classification, structured extraction, summarisation under a length cap, open-ended rewriting, code generation, arithmetic, planning, and a single agent tool-call decision.

The discipline matters more than the coverage. The prompts are frozen in source and may not be edited in place, because a measurement taken in November against an edited prompt is not comparable to one taken in August. The fingerprint of the exact text sent is stored on every row, so an accidental edit shows up as a changed fingerprint rather than as drift. Failed calls are recorded as failures rather than dropped, so a gap in a series always states its reason. And no row can be edited or deleted after it is written.

There is no alerting attached to it yet, deliberately. One reading is not a series, and a comparison written today could not be tested against real history. Shipping the alarm before the readings is how tools end up reporting movement they cannot substantiate. The meter runs now so that the comparison, when it ships, has something real behind it.

What each front actually justifies

  • The market movesProven from our own append-only price ledger. On its own, this is sufficient to make continuous monitoring rational.
  • Your stack maturesReasoned, not measured by us. We will not cite it as our finding until our own history can separate a maturing workload from a growing one.
  • The setup changes silentlyMechanism proven with a captured artifact; the longitudinal case is not yet ours. The monthly meter above is what will settle it.

Notice that the argument does not need all three. If prices move monthly and provably do, then a recommendation set derived from last quarter's prices is a historical document no matter how good it was when written. The other two fronts do not change that conclusion. They change how much of the movement you can currently see, and therefore how much of it is quietly costing you.

What this means for how you buy

None of this makes a one-time audit worthless. An audit finds the standing waste, and standing waste is real money. What it cannot do is hold. The correct mental model is not consulting engagement versus software subscription; it is closer to the difference between having your accounts audited once and having a ledger. The audit tells you where you stood. The ledger tells you where you are.

The test to apply to anyone selling you either is the same test we apply to ourselves in public. Ask which of their claims are measured from their own data, which are mechanisms they can show you an artifact for, and which are reasoning they have not yet proven. A vendor who cannot separate those three about their own product will not separate them about your bill.

The four-rung framework this all builds toward, in one place

Read The CostMyAI Standard

Start with the free level.
See the saving before you pay us anything.