Methodology

How a switch gets proven.

A recommendation you cannot audit is a guess with a dollar sign in front of it. This is the whole decision procedure, written down so it can be checked.

  • 01

    One cost function

    Every price comparison in the product and on this website runs through a single function: input tokens times the input rate plus output tokens times the output rate, on your own observed token mix rather than a list-price headline. There is no second formula anywhere, because two formulas produce two different savings for the same switch.

  • 02

    The quality bar

    A cheaper model is only offered when its third-party benchmark score clears your current model's score minus that evaluation's own measured margin, for that task class. The margin is synced alongside the scores it applies to — never a hardcoded tolerance, and never a number we choose. A separate discrimination check (below) then ensures the difference is large enough to be reliable.

  • 03

    The discrimination guard

    Beyond the margin sits a separate gate: discrimination. If every model's score on a task class falls within a fixed 10-point range on the benchmark's 0–100 scale, the benchmark cannot tell them apart reliably enough to justify any quality claim.

  • 04

    The tie-break

    The cheapest option clearing the bar wins — not the highest-scoring, not a partner's. An exact price tie breaks alphabetically by model, then by host. The rule is fixed so no thumb can rest on the scale.

  • 05

    Refusals

    When nothing clears, you get the refusal and its reason: no baseline price, no baseline score, benchmark not discriminating, nothing cheaper clearing the bar, latency ceiling unmet, or a saving too small to be worth a switch. Refusals are a product surface, not an error state.

  • 06

    Freshness

    Prices and benchmark margins come from public feeds that re-sync continuously. Coverage figures shown publicly are read from the same tables the engine prices against, and are only labelled live once a sync has actually completed. Inside your workspace, every figure carries the timestamp of the run it came from.

  • 07

    What we never hold

    No provider API keys, no prompts, no completions, no user content. The Verification Engine runs in your environment; what reaches us is aggregate metadata only, and the ingest schema rejects anything else.

  • 08

    Retention

    The aggregate usage records pushed by the Verification Engine, and the rollups and recommendations derived from them, are kept for as long as your workspace exists, because a savings figure is only auditable against the history it was computed from. Close your workspace, or ask us to delete it, and that data is removed within 30 days. Public market data — model prices, benchmark scores and their change history — is not customer data and is kept permanently as a public record.

  • 09

    What counts as a price move

    A price move is one observed change to a live listed price for a model on a specific host, recorded with its direction (increase or decrease).

    A model or host appearing for the first time is not a move. A delisting is not a move. A model or host reappearing after being delisted is also not a move. All three are excluded from the count.

    Moves are counted between two of our own pricing syncs, so the number reflects what we actually caught, not what a provider announced. A price that changes and reverts between two syncs is invisible to us.

    The counter covers the current calendar month in UTC and resets on the 1st. The underlying ledger is append-only and permanent — rows cannot be deleted, edited, pruned, or archived — so the window is a read choice, not data loss.

    Coverage started on July 31, 2026.

Every rule above runs on a live catalog.
See it applied to real prices.