Methodology

How a switch gets proven.

A recommendation you cannot audit is a guess with a dollar sign in front of it. This is the whole decision procedure, written down so it can be checked.

  • 01

    One cost function

    Every price comparison in the product and on this website runs through a single function: input tokens times the input rate plus output tokens times the output rate, on your own observed token mix rather than a list-price headline. There is no second formula anywhere, because two formulas produce two different savings for the same switch.

  • 02

    The quality bar

    A cheaper model is only offered when its third-party benchmark score clears your current model's score minus that evaluation's own measured margin, for that task class. The margin is synced alongside the scores it applies to — never a hardcoded tolerance, and never a number we choose.

  • 03

    The discrimination guard

    If every model's score on a task class sits inside the measurement margin, the benchmark cannot tell them apart and must not be used to justify anything. We require the observed spread to exceed twice the margin before a quality claim is possible at all.

  • 04

    The tie-break

    The cheapest option clearing the bar wins — not the highest-scoring, not a partner's. An exact price tie breaks alphabetically by model, then by host. The rule is fixed so no thumb can rest on the scale.

  • 05

    Refusals

    When nothing clears, you get the refusal and its reason: no baseline price, no baseline score, benchmark not discriminating, nothing cheaper clearing the bar, latency ceiling unmet, or a saving too small to be worth a switch. Refusals are a product surface, not an error state.

  • 06

    Freshness

    Prices and benchmark margins come from public feeds that re-sync continuously. Coverage figures shown publicly are read from the same tables the engine prices against, and are only labelled live once a sync has actually completed. Inside your workspace, every figure carries the timestamp of the run it came from.

  • 07

    What we never hold

    No provider API keys, no prompts, no completions, no user content. The Verification Engine runs in your environment; what reaches us is aggregate metadata only, and the ingest schema rejects anything else.

  • 08

    Retention

    The aggregate usage records pushed by the Verification Engine, and the rollups and recommendations derived from them, are kept for as long as your workspace exists, because a savings figure is only auditable against the history it was computed from. Close your workspace, or ask us to delete it, and that data is removed within 30 days. Public market data — model prices, benchmark scores and their change history — is not customer data and is kept permanently as a public record.

Every rule above runs on a live catalog.
See it applied to real prices.