Blog

Why AI costs overshoot the forecast, and what to do in the first week

The gap is rarely a pricing surprise. It is retries, context growth, and workloads nobody re-examined. Here is how to close it.

23 August 2026 · 6 min read

The most common finance complaint about AI is not that it is expensive. It is that it is unpredictable in one direction. Budgets set in good faith are overrun by margins that would be a scandal in any other line item, and the post-mortem almost never finds a price increase. Unit prices generally fell. Volume did something the forecast never modelled.

The five things that actually blow the number

  • Retries and agent loopsAn agentic workflow does not make one call, it makes a chain of them, and every failed step, tool call, and self-correction is billed. A workflow forecast at one call per task can settle at eight in production without anyone changing the code.
  • Context creepPrompts grow. A system prompt gains guardrails, retrieval returns more chunks, conversation history is carried further. Input tokens per call rise quietly for months, and input tokens are the bulk of most bills.
  • Reasoning outputReasoning-style models emit tokens you never see and always pay for. A model swap that looked cost-neutral on the published rate card can double effective cost per task.
  • Unmodelled workloadsThe forecast covered the two workloads that existed at planning time. By month four there are nine, several launched by teams who never touched the budget line.
  • Host driftThe same model is available from several hosts at materially different prices. Traffic lands wherever the first integration pointed, and nobody re-checks after the market moves.

Forecast the driver, not the invoice

A forecast built by extrapolating last quarter's invoice is guessing at the output of a system it does not model. A forecast built on drivers is auditable: calls per unit of business activity, tokens per call split into input and output, and the price per token of the model and host actually serving that workload. When the number moves, you can say which of the three moved and by how much. That is the entire difference between a forecast and a hope.

See where the cheapest host for your model actually is

Open the cheapest API calls report

Track variance weekly, not at month end

Most overruns are visible in week one and discovered in week five. A workload whose tokens per call jumped forty percent after a prompt change announces itself immediately if anyone is watching that ratio, and hides completely inside a monthly total. Weekly variance against the driver forecast, per workload, catches the structural break while it is still cheap to reverse.

The threshold matters as much as the cadence. An alert on every fluctuation gets muted within a fortnight. An alert on a sustained change in tokens per call, or on a new workload appearing with no forecast line, stays credible because it fires rarely and is right when it does.

Budget a range, and say why

A single-point AI budget is a fiction with a decimal place. The honest artefact is a range with a stated basis: the floor assumes current usage patterns hold, the mid case assumes planned launches ship on schedule, the ceiling assumes usage per user rises at the rate it has been rising. Finance can plan against a range with reasoning behind it. Nobody can plan against a confident number that turns out to be wrong by half.

What the market actually did this month

Read the latest intelligence

The first week of work

Split current spend by workload rather than by provider. For the top three, record calls per task and tokens per call as they stand today, because that is the baseline every future variance is measured against. Check whether each one is served by the cheapest host offering the identical model. Then set a weekly review on those two ratios. None of this requires new tooling, and it converts the overshoot from an annual surprise into a weekly, fixable signal.

Start with the free level.
See the saving before you pay us anything.