Report
The cheapest API call, model by model.
Every model we hold a verified price for, ranked by what a million input tokens costs at its cheapest serving provider — and how much more the dearest one charges for the exact same weights.
Cheapest verified price per model
Prices in USD per million tokens, cheapest first. Each row shows the cheapest verified host for input price and what the same host charges for output tokens.
DeepSeek V4 Flash 0731
deepseek
Relace
$0.01
$1.28
−97% cheaper on input price
Granite 4.0 Micro
ibm-granite
Cloudflare
$0.02
$0.11
—
DeepSeek V4 Flash 0423
deepseek
Relace
$0.02
$1.28
−96% cheaper on input price
gpt-oss-20b
openai
Darkbloom
$0.02
$0.09
−76% cheaper on input price
Mistral Nemo
mistralai
DekaLLM
$0.02
$0.03
−88% cheaper on input price
DeepSeek V4.1 Flash
deepseek
Relace
$0.02
$1.20
−94% cheaper on input price
Llama 3.1 8B Instruct
meta-llama
DeepInfra
$0.02
$0.04
−91% cheaper on input price
Ling 3.0 Flash
inclusionai
Novita
$0.02
$0.06
−65% cheaper on input price
Ling 3.0 Flash VL
inclusionai
Novita
$0.02
$0.06
−65% cheaper on input price
Qwen3.8 27B
qwen
Wafer
$0.02
$4.35
−98% cheaper on input price
gpt-oss-20b (batch)
openai
DeepInfra
$0.02
$0.11
—
GPT-5 Nano
openai
OpenAI
$0.03
$0.20
−50% cheaper on input price
GPT-5 Nano (batch)
openai
OpenAI
$0.03
$0.20
—
Nex-N2.5-Mini
nex-agi
Nex AGI
$0.03
$0.10
—
Llama 3.2 1B Instruct
meta-llama
Cloudflare
$0.03
$0.20
—
gpt-oss-120b (batch)
openai
DeepInfra
$0.03
$0.14
—
gpt-oss-120b
openai
CoreWeave
$0.03
$0.17
−91% cheaper on input price
Qwen3.7 Flash
qwen
Alibaba
$0.03
$0.13
—
Schematron V2 Turbo
inference-net
InferenceNet
$0.03
$0.15
—
GLM 5.3 Flash
z-ai
OpenInference
$0.03
$0.66
−89% cheaper on input price
Nova Micro 1.0
amazon
Amazon Bedrock
$0.04
$0.14
—
Command R7B (12-2024)
cohere
Cohere
$0.04
$0.15
—
Nemotron 3.5 Lightning
nvidia
Darkbloom
$0.04
$0.18
−44% cheaper on input price
GLM 5.3
z-ai
Relace
$0.04
$12.00
−97% cheaper on input price
Llama 3 8B Lunaris
sao10k
DeepInfra
$0.04
$0.05
−20% cheaper on input price
Mercury 2.5
inception
Inception
$0.04
$0.15
—
Mercury 2.5 Preview
inception
Inception
$0.04
$0.15
—
Gemma 4 26B A4B
Darkbloom
$0.04
$0.22
−72% cheaper on input price
Ling 3.0 Flash Fin
inclusionai
Novita
$0.04
$0.12
−30% cheaper on input price
Hy-MT2-1.8B
tencent
Tencent
$0.04
$0.18
—
Qwen3 30B A3B Instruct 2507
qwen
StreamLake
$0.05
$0.19
−63% cheaper on input price
Claude Haiku 5.5 (batch)
anthropic
Anthropic
$0.05
$0.25
—
Gemini 2.5 Flash Lite
Google AI Studio
$0.05
$0.20
−50% cheaper on input price
Gemini 2.5 Flash Lite (batch)
$0.05
$0.20
—
Gemma 3 12B
DeepInfra
$0.05
$0.15
—
Gemma 3 4B
DeepInfra
$0.05
$0.10
—
GPT-4.1 Nano (batch)
openai
OpenAI
$0.05
$0.20
—
GPT-6 Luna
openai
OpenAI
$0.05
$0.25
−55% cheaper on input price
GPT-6 Luna (batch)
openai
OpenAI
$0.05
$0.25
—
GPT-6 Luna Pro
openai
OpenAI
$0.05
$0.25
−50% cheaper on input price
Showing the 40 cheapest of 480 tracked models. Browse the full catalog to filter by vendor, quality score or host spread.
Which provider is cheapest, host by host
For every model sold by more than one provider, exactly one host has the lowest input rate. This counts those wins. A provider that carries many models but wins few is convenient rather than cheap.
DeepInfra
78
2735%
$0.14
OpenAI
86
2327%
$1.00
Amazon Bedrock
39
1436%
$2.20
52
1121%
$0.38
Azure
56
1120%
$2.00
Alibaba
59
1017%
$0.30
Novita
71
811%
$0.30
Darkbloom
9
778%
$0.05
Relace
11
764%
$0.17
GMICloud
26
623%
$0.30
Google AI Studio
16
531%
$0.38
StreamLake
22
418%
$0.34
Parasail
38
38%
$0.15
SiliconFlow
40
38%
$0.27
Minimax
8
338%
$0.30
DekaLLM
10
220%
$0.08
Wafer
9
222%
$0.10
CoreWeave
19
211%
$0.22
Venice
37
25%
$0.33
Inceptron
6
233%
$0.60
Want this priced against your own volumes? Use the LLM pricing comparison calculator.
How to read a cheap API call
A price is per token, not per call. An API call costs input tokens plus output tokens. Two providers can advertise the same headline rate and still bill differently once the output side is counted, which is why the blended column exists above.
The same model has more than one price. Open weight models are sold by several serving providers at once. Where that gap is wide, the cheapest call is the same model on a different host, not a weaker model.
Cheapest is not automatically correct. A model that is cheap per token but needs two attempts, or emits three times the output, is not cheap per finished task. We publish the benchmark scores next to every price on the model catalog so a quality claim has something to clear.
Prices come from the same catalog our engine prices against, and aggregator listings are excluded from provider-to-provider gaps. The full sourcing rules are in the methodology.
Price it against your own traffic.
Compare reads this same catalog and shows what your current models would cost on the cheapest verified host. Free, forever.