Guide

AI cost management,
without guessing.

Most teams discover their AI spend the same way: an invoice arrives, nobody can say which feature caused it, and the only lever anyone trusts is using the model less. This guide sets out what actually drives the bill and the order in which to attack it.

What actually drives an AI bill

01

Volume you cannot attribute

A provider invoice is one number for an entire company. Without per-workload attribution, nobody can say whether support summaries or the internal coding assistant caused the increase, so nobody owns the fix.

02

Output length, not input length

Output tokens usually cost several times more than input tokens. A prompt change that makes answers longer raises the bill more than the same change applied to context.

03

The wrong host for the right model

Identical open-weight models are served by multiple providers at different rates. Paying the dearest host for the same weights is pure waste, and it is invisible on an invoice.

04

Prices that move under you

Published rates change without announcements. A route that was cheapest last quarter may not be cheapest today, and nothing tells you when it stopped being true.

05

Retries, reasoning and cache misses

Failed calls, reasoning tokens and lost prompt-cache hits are all billed. They rarely appear in a cost model built from a pricing page.

The five steps, in the order that pays

The sequence matters more than the tooling. Teams that start at step three spend weeks evaluating models while the same-model waste from step two sits untouched.

Step 1

Get the spend visible per workload

Break the invoice into the jobs that produced it. Until a cost has an owner and a purpose, every reduction argument is a matter of opinion.

Step 2

Reprice the same work at every host

For each workload, compute what the identical calls would have cost at every provider serving that model. This is the only saving with no quality question attached.

Step 3

Prove quality before changing a model

Where a different model is genuinely cheaper, it has to pass your own task before it earns the traffic. A cheaper model that needs a retry or writes three times as much is not cheaper.

Step 4

Right-size the model to the task

Most workloads are running on a model larger than the job needs. Match the tier to the difficulty of the task rather than to the most capable model available.

Step 5

Keep watching, because prices move

Cost management is not a project with an end date. Rates change, new hosts appear, and traffic shifts. Whatever you decide today needs re-checking against the next price move.

What we track so you do not have to

390

Models priced

71

Providers with verified live prices

2,238

Price moves caught this month

Every rate is a published list price we hold a verified record for. Sourcing rules are in the methodology, the full catalog is in Models, and monthly price movement is published in Intelligence.

Common questions

What is AI cost management?

AI cost management is the practice of measuring what every AI workload costs, attributing that cost to the team or feature that caused it, and then reducing it without degrading output quality. It differs from ordinary cloud cost management because the unit of spend is a token, the price of that token changes without notice, and two providers can serve identical model weights at very different rates.

Why does an AI bill rise while token prices fall?

Volume and verbosity grow faster than published rates fall. More features call a model, prompts get longer, retries and reasoning tokens accumulate, and output length drifts upward. A falling per-token price applied to a rising token count still produces a larger invoice.

What is the cheapest way to cut AI spend without a quality risk?

Move the same model to the cheapest verified host serving it. The weights are identical, so output does not change, and there is nothing to re-evaluate. Only after that is exhausted does swapping to a different model become worth the evaluation cost.

Do AI cost tools need your provider API keys?

No, and they should not have them. Cost analysis needs usage and price records, not the ability to make calls on your behalf. CostMyAI never holds your provider keys.

More answers in the full FAQ.

Start with step two. It costs nothing.

Compare reads your real usage and names the cheaper route for every workload, with no model change and no key handover. Free, forever.