Report

The cheapest API call, model by model.

Every model we hold a verified price for, ranked by what a million input tokens costs at its cheapest serving provider — and how much more the dearest one charges for the exact same weights.

413Models with a verified price
76Serving providers
$0.01Cheapest rate / 1M in
98%Widest same-model gap (input price)

Cheapest verified price per model

Prices in USD per million tokens, cheapest first. Each row shows the cheapest verified host for input price and what the same host charges for output tokens.

DeepSeek V4 Flash 0731

deepseek

Relace

$0.01

$1.28

−97% cheaper on input price

Granite 4.0 Micro

ibm-granite

Cloudflare

$0.02

$0.11

—

DeepSeek V4 Flash 0423

deepseek

Relace

$0.02

$1.28

−96% cheaper on input price

gpt-oss-20b

openai

Darkbloom

$0.02

$0.09

−76% cheaper on input price

Mistral Nemo

mistralai

DekaLLM

$0.02

$0.03

−88% cheaper on input price

DeepSeek V4.1 Flash

deepseek

Relace

$0.02

$1.20

−94% cheaper on input price

Llama 3.1 8B Instruct

meta-llama

DeepInfra

$0.02

$0.04

−91% cheaper on input price

Ling 3.0 Flash

inclusionai

Novita

$0.02

$0.06

−65% cheaper on input price

Ling 3.0 Flash VL

inclusionai

Novita

$0.02

$0.06

−65% cheaper on input price

Qwen3.8 27B

qwen

Wafer

$0.02

$4.35

−98% cheaper on input price

gpt-oss-20b (batch)

openai

DeepInfra

$0.02

$0.11

—

GPT-5 Nano

openai

OpenAI

$0.03

$0.20

−50% cheaper on input price

GPT-5 Nano (batch)

openai

OpenAI

$0.03

$0.20

—

Nex-N2.5-Mini

nex-agi

Nex AGI

$0.03

$0.10

—

Llama 3.2 1B Instruct

meta-llama

Cloudflare

$0.03

$0.20

—

gpt-oss-120b (batch)

openai

DeepInfra

$0.03

$0.14

—

gpt-oss-120b

openai

CoreWeave

$0.03

$0.17

−91% cheaper on input price

Qwen3.7 Flash

qwen

Alibaba

$0.03

$0.13

—

Schematron V2 Turbo

inference-net

InferenceNet

$0.03

$0.15

—

GLM 5.3 Flash

z-ai

OpenInference

$0.03

$0.66

−89% cheaper on input price

Nova Micro 1.0

amazon

Amazon Bedrock

$0.04

$0.14

—

Command R7B (12-2024)

cohere

Cohere

$0.04

$0.15

—

Nemotron 3.5 Lightning

nvidia

Darkbloom

$0.04

$0.18

−44% cheaper on input price

GLM 5.3

z-ai

Relace

$0.04

$12.00

−97% cheaper on input price

Llama 3 8B Lunaris

sao10k

DeepInfra

$0.04

$0.05

−20% cheaper on input price

Mercury 2.5

inception

Inception

$0.04

$0.15

—

Mercury 2.5 Preview

inception

Inception

$0.04

$0.15

—

Gemma 4 26B A4B

google

Darkbloom

$0.04

$0.22

−72% cheaper on input price

Ling 3.0 Flash Fin

inclusionai

Novita

$0.04

$0.12

−30% cheaper on input price

Hy-MT2-1.8B

tencent

Tencent

$0.04

$0.18

—

Qwen3 30B A3B Instruct 2507

qwen

StreamLake

$0.05

$0.19

−63% cheaper on input price

Claude Haiku 5.5 (batch)

anthropic

Anthropic

$0.05

$0.25

—

Gemini 2.5 Flash Lite

google

Google AI Studio

$0.05

$0.20

−50% cheaper on input price

Gemini 2.5 Flash Lite (batch)

google

Google

$0.05

$0.20

—

Gemma 3 12B

google

DeepInfra

$0.05

$0.15

—

Gemma 3 4B

google

DeepInfra

$0.05

$0.10

—

GPT-4.1 Nano (batch)

openai

OpenAI

$0.05

$0.20

—

GPT-6 Luna

openai

OpenAI

$0.05

$0.25

−55% cheaper on input price

GPT-6 Luna (batch)

openai

OpenAI

$0.05

$0.25

—

GPT-6 Luna Pro

openai

OpenAI

$0.05

$0.25

−50% cheaper on input price

Showing the 40 cheapest of 480 tracked models. Browse the full catalog to filter by vendor, quality score or host spread.

Which provider is cheapest, host by host

For every model sold by more than one provider, exactly one host has the lowest input rate. This counts those wins. A provider that carries many models but wins few is convenient rather than cheap.

DeepInfra

78

2735%

$0.14

OpenAI

86

2327%

$1.00

Amazon Bedrock

39

1436%

$2.20

Google

52

1121%

$0.38

Azure

56

1120%

$2.00

Alibaba

59

1017%

$0.30

Novita

71

811%

$0.30

Darkbloom

9

778%

$0.05

Relace

11

764%

$0.17

GMICloud

26

623%

$0.30

Google AI Studio

16

531%

$0.38

StreamLake

22

418%

$0.34

Parasail

38

38%

$0.15

SiliconFlow

40

38%

$0.27

Minimax

8

338%

$0.30

DekaLLM

10

220%

$0.08

Wafer

9

222%

$0.10

CoreWeave

19

211%

$0.22

Venice

37

25%

$0.33

Inceptron

6

233%

$0.60

Want this priced against your own volumes? Use the LLM pricing comparison calculator.

How to read a cheap API call

A price is per token, not per call. An API call costs input tokens plus output tokens. Two providers can advertise the same headline rate and still bill differently once the output side is counted, which is why the blended column exists above.

The same model has more than one price. Open weight models are sold by several serving providers at once. Where that gap is wide, the cheapest call is the same model on a different host, not a weaker model.

Cheapest is not automatically correct. A model that is cheap per token but needs two attempts, or emits three times the output, is not cheap per finished task. We publish the benchmark scores next to every price on the model catalog so a quality claim has something to clear.

Prices come from the same catalog our engine prices against, and aggregator listings are excluded from provider-to-provider gaps. The full sourcing rules are in the methodology.

Price it against your own traffic.

Compare reads this same catalog and shows what your current models would cost on the cheapest verified host. Free, forever.