How It Works
Connect once. Governed decisions on every workload.
Run the CostMyAI Verification Engine in your environment, point your SDK base URL at it, and keep your provider keys exactly where they are today. CostMyAI sees only metadata.
Connect
Point your application's SDK base URL at the Verification Engine that runs in your own environment. Your code, your provider keys, your contracts — nothing moves to us.
Environment
OPENAI_BASE_URL=http://localhost:8787/v1SDK
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "http://localhost:8787/v1" }); // apiKey unchangedYour application keeps sending its own provider key. Only the base URL changes.
- The Verification Engine runs as a small container in your environment. We never hold your provider keys.
- Nothing to migrate and nothing to rewrite: the request and response pass through unchanged.
- By default the engine does not read request bodies. What reaches CostMyAI is token counts, model names and request counts.
- If you turn on local classification, request bodies are read inside your own environment and only the task label is sent to us.
- Only aggregate metadata leaves your environment, on its own path, after the response is served.
Map
We read your real spend, group it by workload, and benchmark every model against the live catalog, which re-syncs continuously, so a verdict is always measured against today's prices, not last quarter's. The buy-side view, not the vendor's.
- Traffic is grouped by workload rather than by raw model name, so a verdict covers the job you are actually running.
- Prices come from the tracked provider feeds and are re-synced continuously, with every host priced separately.
- Scores come from published evaluations, carrying the suite, the task class and the measurement margin they were taken with.
Verdict
See which switches hold quality on real benchmarks, and which ones we refuse to certify. A governed decision names what it cannot prove.
- Every recommendation states the measurement it rests on: the suite, the score and the margin.
- Where nothing clears the bar, you get a refusal with the reason instead of a switch.
- Same-model, cheaper-host moves are separated from quality-matched model changes, because they carry different risk.
Switch
Switch the workloads that hold quality. Keep the savings. Leave the rest exactly where they are. Not paying more than you need to, on the record and defensible.
- Switches are activated per workload, can be paused, and roll back in one click.
- On Govern the same decision is re-checked at the moment of action, not only at the moment of evaluation.
- Every automated decision is written to an audit trail you can hand to finance.
The request path
Runs in your environment. Sees only metadata.
The Verification Engine sits in your stack as middleware. Requests pass through unchanged; only token counts and model names leave your environment by default. No prompt content leaves your environment unless you explicitly opt in to remote classification — see how that opt-in works.
01request
Your App
Makes API requests
02forwarded unchanged
Verification Engine
Middleware in your environment
03provider
AI Provider
OpenAI, Anthropic, Gemini, others
03 · RESPONSE
Travels the same path in reverse — provider → engine → your app. The engine counts tokens; it does not read the body.
04 · METADATA ONLY
The only leg that leaves your environment carries nothing we could reconstruct a prompt from.
Aggregate rows. Token counts, model names, hosts. No prompt content by default.
Level by level
What each level actually does.
Same four steps at every level. What changes is how far the verdict is allowed to go.
Compare
You want to know if a cheaper host exists for the same model.
Certify
You need to prove quality before you switch to a cheaper model.
Rightsize
You want the switch executed, but you stay in control.
Govern
You want switching continuous, bounded, and audited.
Level 1
Compare
Same model, run through whichever provider charges less for it. Nothing about the output changes. Only who gets paid.
Free forever
- Same model, cheaper host — across every priced host
- Live price catalog, re-synced continuously across every tracked provider
- Metadata-only ingest, no provider keys
- Spend, tokens and requests over 24h / 7d / 30d
- Up to 3 workspace members

Level 2
Certify
A cheaper model, proven to score the same as what you're running today. 'Certified' means it cleared a benchmark test built to catch the difference. Not just a claim that it's just as good.
From $58/mo billed yearly
- Everything in Compare
- Quality-matched cheaper models, cheapest that clears the bar
- Published evaluation, score and measurement margin per claim
- Refusals with reasons when nothing clears
- Objectives: cost, latency ceiling, quality floor
- Invoice reconciliation against your own provider bills
- Up to 10 workspace members

Level 3
Rightsize
Matches the model to what the task actually requires. Some tasks are running on far more model than the work needs. Rightsize points those at a model built for that size of problem, not a smaller budget.
From $324/mo billed yearly
- Everything in Certify
- Oversized-workload detection per task class
- Manual switch activation, pause and one-click rollback
- Up to 25 workspace members

Level 4
Govern
Everything Compare, Certify, and Rightsize already proved safe, applied automatically. Without you clicking anything. Every switch it runs on its own already cleared the same evidence bar Compare, Certify, and Rightsize use for you.
From $749/mo billed yearly
- Everything in Rightsize
- Autonomous switching inside the equivalence band
- Re-checked at the moment of action, not only at evaluation
- Full audit trail of every automated decision
- Unlimited workspace members

What teams ask before they start
Security, quality, and fallback.
How do I know if switching to a cheaper model will hurt my output quality?
Test it against an independent, task specific benchmark before you switch any real traffic, not a general leaderboard score and not a spot check on a handful of examples. The comparison needs to account for the benchmark's own measurement uncertainty too. Two models scoring within a few points of each other might not be meaningfully different at all, or they might be, and you cannot tell which without knowing the margin.
Is it risky to route AI traffic across multiple providers?
The bigger risk is usually the opposite: depending entirely on one provider. A single vendor relationship means no fallback if that provider raises prices, changes availability, or has an outage, and no leverage to negotiate anything once you are fully dependent. A properly built multi-provider setup, where switching is a certified, proven decision rather than a guess, reduces risk rather than adding it.
What happens if no safe alternative exists for a given model?
The honest answer should be that nothing gets switched, with the specific reason stated, not a quieter downgrade suggested anyway. If a cheaper candidate cannot be shown to perform equivalently on an independent benchmark for that exact task, recommending it anyway trades a visible cost saving for an invisible quality risk, which is not actually a saving.
Does CostMyAI ever hold or see my provider API keys?
No. The component that reads your usage data runs inside your own environment, not ours, and your provider keys never leave it. We built the architecture this way specifically so that getting visibility into your AI spend never requires a new leap of trust in a third party holding your credentials.
What data does CostMyAI actually see, if not my API keys?
Aggregate, provider neutral records of usage and billed spend, pushed from your own environment. Not your prompts, not your model outputs, unless you explicitly choose to share that separately.
Does the Verification Engine add latency or break streaming?
The Verification Engine forwards requests and streams responses without buffering them. It counts tokens from the response headers or tail, so the added latency is typically sub-millisecond for non-streaming calls and zero for streaming bodies because they pass through unchanged. If the Verification Engine stops, your application falls back to the original provider base URL instantly — there is no dependency on CostMyAI being up for your traffic to keep flowing.
What happens if CostMyAI is down when a switch is supposed to run?
On Rightsize and below, switches are manual — you review and activate them yourself. On Govern, the Verification Engine re-checks the decision at the moment of action using a cached policy. If the policy cannot be refreshed and the decision is no longer within the configured equivalence band, the switch is refused rather than executed. The default is always to do nothing rather than guess.
Keep your keys. Change one URL. Then the evidence.
Compare is free forever, and it needs no provider keys to start.