How It Works
Connect once. Governed decisions on every workload.
Four steps, no manual exports — from one environment variable to a benchmark-backed verdict on every workload, and a switch you can defend afterwards.
Connect
Point your application at the Verification Engine endpoint. One environment variable change. Requests forward to your provider unchanged. What reaches us is token counts, model names, and request counts. Never your prompt content: by default the engine does not read request bodies at all, and if you turn on local classification it reads them inside your own environment and sends us only the task label.
- The engine runs as middleware inside your own environment. We never hold your provider keys.
- Nothing to migrate and nothing to rewrite: the request and the response pass through unchanged.
- Only aggregate metadata leaves your environment, on its own path, after the response is served.
Map
We read your real spend, group it by workload, and benchmark every model against the live catalog — which re-syncs continuously, so a verdict is always measured against today's prices, not last quarter's. The buy-side view, not the vendor's.
- Traffic is grouped by workload rather than by raw model name, so a verdict covers the job you are actually running.
- Prices come from the tracked provider feeds and are re-synced continuously, with every host priced separately.
- Scores come from published evaluations, carrying the suite, the task class and the measurement margin they were taken with.
Verdict
See which switches hold quality on real benchmarks, and which ones we refuse to certify. A governed decision names what it cannot prove.
- Every recommendation states the measurement it rests on: the suite, the score and the margin.
- Where nothing clears the bar, you get a refusal with the reason instead of a switch.
- Same-model, cheaper-host moves are separated from quality-matched model changes, because they carry different risk.
Switch
Switch the workloads that hold quality. Keep the savings. Leave the rest exactly where they are. Not paying more than you need to, on the record and defensible.
- Switches are activated per workload, can be paused, and roll back in one click.
- On Govern the same decision is re-checked at the moment of action, not only at the moment of evaluation.
- Every automated decision is written to an audit trail you can hand to finance.
The request path
Runs in your environment. Sees only metadata.
The Verification Engine sits in your stack as middleware. Requests pass through unchanged; only token counts and model names leave your environment. Never prompt content.
01request
Your App
Makes API requests
02forwarded unchanged
Verification Engine
Middleware in your environment
03provider
AI Provider
OpenAI, Anthropic, Gemini, others
03 · RESPONSE
Travels the same path in reverse — provider → engine → your app. The engine counts tokens; it does not read the body.
04 · METADATA ONLY
The only leg that leaves your environment carries nothing we could reconstruct a prompt from.
Aggregate rows. Token counts, model names, hosts. No prompt content, ever.
Level by level
What each level actually does.
Same four steps at every level. What changes is how far the verdict is allowed to go.
Level 1
Compare
Same model, cheaper host.
Free forever
- Same model, cheaper host — across every priced host
- Live price catalog, re-synced continuously across every tracked provider
- Metadata-only ingest, no provider keys
- Spend, tokens and requests over 24h / 7d / 30d
- Up to 3 workspace members

Level 2
Certify
Plus quality-matched cheaper models.
From $58/mo billed yearly
- Everything in Compare
- Quality-matched cheaper models, cheapest that clears the bar
- Published evaluation, score and measurement margin per claim
- Refusals with reasons when nothing clears
- Objectives: cost, latency ceiling, quality floor
- Invoice reconciliation against your own provider bills
- Up to 10 workspace members

Level 3
Rightsize
Plus oversized-workload detection and manual switching.
From $324/mo billed yearly
- Everything in Certify
- Oversized-workload detection per task class
- Manual switch activation, pause and one-click rollback
- Up to 25 workspace members

Level 4
Govern
Plus autonomous switching by CostMyAI.
From $749/mo billed yearly
- Everything in Rightsize
- Autonomous switching inside the equivalence band
- Re-checked at the moment of action, not only at evaluation
- Full audit trail of every automated decision
- Unlimited workspace members

One environment variable. Then the evidence.
Compare is free forever, and it needs no provider keys to start.