Blog

Why your AI bill keeps rising even as token prices fall

Token prices have collapsed since 2023. Bills have not. The variable that actually drives your invoice is consumption.

18 July 2026 · 5 min read

If your AI bill has gone up this year, you are not imagining it, and you are also not wrong that model pricing has been falling. Both things are true at once, and understanding why is the key to actually controlling AI spend rather than just watching it.

The headline number everyone quotes

Per token pricing for frontier grade AI models has fallen dramatically since early 2023, by some measures roughly ninety percent or more for comparable capability. New model generations routinely launch cheaper than the ones they replace, and providers compete aggressively on price as open weight alternatives put pressure on every tier of the market. On paper, this should be the best possible news for anyone budgeting AI spend.

The number that explains why bills are not falling anyway

Usage has grown even faster than price has fallen. Reporting drawing on real enterprise and developer token volume found business token consumption growing roughly ten times faster than token spend over a recent twelve month period, meaning the industry is using vastly more tokens per dollar than before, and still spending more overall dollars than before. The math only resolves one way: consumption is the variable actually driving the bill, not price.

Agentic workloads are the accelerant

The single biggest reason consumption has exploded is the shift from simple chatbot style requests to agentic workflows. An agent that plans, calls tools, checks its own work, and iterates can consume ten to thirty times more tokens per completed task than a single chat completion, because every one of those internal steps is itself a model call that gets billed. A cheaper price per token does very little to offset a workload that is quietly generating thirty times as many tokens to do the same job.

This is not a hypothetical. Public reporting has documented enterprise engineering teams exhausting an entire year's AI coding budget in a matter of months once agentic coding tools became the default way of working, purely on volume, with per engineer costs running well into four figures a month.

See real, live price moves across every model you use

View the Intelligence page

Why this matters for how you budget

The practical implication is that tracking price alone tells you almost nothing useful about where your budget is going. A team that only watches published per token rates will be blindsided by a bill that keeps climbing even as every rate card they can find keeps getting cheaper. The number that actually predicts your bill is consumption, broken down by workload, by model, and ideally by the specific feature or team driving it.

What to actually do about it

The lever that works is not negotiating a better rate, it is understanding which workloads are consuming disproportionately and whether they need to. Some tasks genuinely require a frontier model's reasoning depth. A large share of production traffic, classification, extraction, routing, simple summarization, does not, and running it on an oversized model is a cost decision nobody made deliberately, it just accumulated.

Seeing this requires a live, accurate view of actual spend by model and by host, not a monthly invoice summary and not a token count multiplied by a list price that may not reflect volume discounts, caching, or batch processing already in effect. The organizations getting ahead of rising bills are the ones treating consumption visibility as a first class metric, not an afterthought to the price they thought they negotiated.

Start with the free level.
See the saving before you pay us anything.