Report
The cheapest API call, model by model.
Every model we hold a verified price for, ranked by what a million input tokens costs at its cheapest serving provider — and how much more the dearest one charges for the exact same weights.
Cheapest verified price per model
Prices in USD per million tokens, cheapest first. The blended column is one million input plus one million output at the same host, which is the number most teams actually feel on an invoice.
Ling-2.6-flash
inclusionai
Novita
$0.01
$0.03
$0.04
Granite 4.0 Micro
ibm-granite
Cloudflare
$0.02
$0.11
$0.13
Mistral Nemo
mistralai
DeepInfra
$0.02
$0.03
$0.05−132%
Llama 3.1 8B Instruct
meta-llama
DeepInfra
$0.02
$0.04
$0.06−1000%
Ling-3.0-flash
inclusionai
Novita
$0.02
$0.06
$0.08−186%
GPT-5 Nano
openai
OpenAI
$0.03
$0.20
$0.23−100%
GPT-5 Nano (batch)
openai
OpenAI
$0.03
$0.20
$0.23
Nex-N2-Mini
nex-agi
Nex AGI
$0.03
$0.10
$0.13
Llama 3.2 1B Instruct
meta-llama
Cloudflare
$0.03
$0.20
$0.23
gpt-oss-120b
openai
CoreWeave
$0.03
$0.17
$0.20−1067%
gpt-oss-20b
openai
CoreWeave
$0.03
$0.13
$0.16−150%
Qwen3.7 Flash
qwen
Alibaba
$0.03
$0.13
$0.16
Solar Pro 4
upstage
Upstage
$0.03
$0.12
$0.15
Nova Micro 1.0
amazon
Amazon Bedrock
$0.04
$0.14
$0.18
Command R7B (12-2024)
cohere
Cohere
$0.04
$0.15
$0.19
DeepSeek V4 Flash 0731
deepseek
OpenInference
$0.04
$0.13
$0.12−1000%
Llama 3 8B Lunaris
sao10k
DeepInfra
$0.04
$0.05
$0.09−25%
Gemma 4 26B A4B
Darkbloom
$0.04
$0.22
$0.26−257%
Hy-MT2-1.8B
tencent
Tencent
$0.04
$0.18
$0.22
Qwen3 30B A3B Instruct 2507
qwen
StreamLake
$0.05
$0.19
$0.24−170%
DeepSeek V4 Flash 0423
deepseek
StreamLake
$0.05
$0.10
$0.15−801%
Gemini 2.5 Flash Lite
Google AI Studio
$0.05
$0.20
$0.25−100%
Gemini 2.5 Flash Lite (batch)
$0.05
$0.20
$0.25
Gemma 3 12B
DeepInfra
$0.05
$0.15
$0.20
Gemma 3 4B
DeepInfra
$0.05
$0.10
$0.15
GPT-4.1 Nano (batch)
openai
OpenAI
$0.05
$0.20
$0.25
GPT-5.4 Nano (batch)
openai
OpenAI
$0.05
$0.31
$0.36
GPT-5.6 Luna (batch)
openai
OpenAI
$0.05
$0.30
$0.35
GPT-5.6 Luna Pro (batch)
openai
OpenAI
$0.05
$0.30
$0.35
Granite 4.1 8B
ibm-granite
CoreWeave
$0.05
$0.10
$0.15
Llama 3.2 3B Instruct
meta-llama
Parasail
$0.05
$0.33
$0.38−2%
Mistral Small 3
mistralai
DeepInfra
$0.05
$0.08
$0.13
Nemotron 3 Nano 30B A3B
nvidia
Crusoe
$0.05
$0.20
$0.25
Gemma 3n 4B
Together
$0.06
$0.12
$0.18
GLM 4.7 Flash
z-ai
DeepInfra
$0.06
$0.40
$0.46−17%
Laguna XS 2.1
poolside
Poolside
$0.06
$0.12
$0.18
MythoMax 13B
gryphe
NextBit
$0.06
$0.06
$0.12−567%
Nova Lite 1.0
amazon
Amazon Bedrock
$0.06
$0.24
$0.30
GPT-5 Mini (batch)
openai
OpenAI
$0.06
$0.50
$0.56
Qwen3.5-Flash
qwen
Alibaba
$0.07
$0.26
$0.33
Showing the 40 cheapest of 388 tracked models. Browse the full catalog to filter by vendor, quality score or host spread.
Which provider is cheapest, host by host
For every model sold by more than one provider, exactly one host has the lowest input rate. This counts those wins. A provider that carries many models but wins few is convenient rather than cheap.
DeepInfra
73
3244%
$0.23
OpenAI
79
1519%
$1.00
50
1224%
$0.60
Amazon Bedrock
30
1240%
$2.20
Alibaba
53
1121%
$0.26
Azure
44
1125%
$2.00
Novita
70
69%
$0.30
StreamLake
24
625%
$0.34
GMICloud
18
528%
$0.29
Parasail
40
38%
$0.21
CoreWeave
21
314%
$0.25
Minimax
8
338%
$0.30
Baidu
9
333%
$0.35
NextBit
7
343%
$0.40
Darkbloom
2
2100%
$0.07
OpenInference
2
2100%
$0.08
AtlasCloud
27
27%
$0.30
Google AI Studio
14
214%
$0.38
AkashML
5
120%
$0.14
Nebius
11
19%
$0.15
Want this priced against your own volumes? Use the LLM pricing comparison calculator.
How to read a cheap API call
A price is per token, not per call. An API call costs input tokens plus output tokens. Two providers can advertise the same headline rate and still bill differently once the output side is counted, which is why the blended column exists above.
The same model has more than one price. Open weight models are sold by several serving providers at once. Where that gap is wide, the cheapest call is the same model on a different host, not a weaker model.
Cheapest is not automatically correct. A model that is cheap per token but needs two attempts, or emits three times the output, is not cheap per finished task. We publish the benchmark scores next to every price on the model catalog so a quality claim has something to clear.
Prices come from the same catalog our engine prices against, and aggregator listings are excluded from provider-to-provider gaps. The full sourcing rules are in the methodology.
Price it against your own traffic.
Compare reads this same catalog and shows what your current models would cost on the cheapest verified host. Free, forever.