Use case · Reasoning

Cheapest reasoning LLMs

The cheapest models that hold up on GPQA, AIME, MATH, MMLU, HLE. Ranked by price per 1M input tokens.

Models30
Cheapest$0.00
ScopeGPQA · AIME · MATH
What this page is
This page ranks every model with credible reasoning scores (GPQA, AIME, MATH, MMLU, HLE, DROP, BBH) by input price. Reasoning models burn a lot of thinking tokens, so the headline input price is only part of the bill. The cheap end is dominated by open-source reasoners like DeepSeek, Qwen3, and GLM. Premium o-series and Claude Opus sit at the top of the price scale. Pair with our cost calculator to model real workloads.

Models with credible reasoning scores, cheapest first.

#ModelIn $/1MOut $/1MType
1Google DeepMind logoGemma 3 27B (free)$0.00$0.00OSS
2OpenAI logogpt-oss-120b (free)$0.00$0.00OSS
3OpenAI logogpt-oss-20b (free)$0.00$0.00OSS
4liquid logoLFM2.5-2.6B (free)$0.00$0.00Closed
5Meta logoLlama 3.2 3B Instruct (free)$0.00$0.00OSS
6Meta logoLlama 3.3 70B Instruct (free)$0.00$0.00OSS
7Cohere logoNorth Mini Code (free)$0.00$0.00OSS
8DeepSeek logoDeepSeek V4 Flash 0731$0.02$1.28Closed
9ibm-granite logoGranite 4.0 Micro$0.02$0.11OSS
10OpenAI logogpt-oss-20b$0.02$0.09OSS
11Mistral AI logoMistral Nemo$0.02$0.03OSS
12Meta logoLlama 3.1 8B Instruct$0.02$0.03OSS
13DeepSeek logoDeepSeek V4 Flash$0.03$1.28OSS
14Meta logoLlama 3.2 1B Instruct$0.03$0.20OSS
15Google DeepMind logoGemma 2 9B$0.03$0.09OSS
16Alibaba Qwen logoQwen2.5 Coder 7B Instruct$0.03$0.09OSS
17Alibaba Qwen logoQwen3.7 Flash$0.03$0.13Closed
18OpenAI logogpt-oss-120b$0.04$0.17OSS
19Alibaba Qwen logoQwen3 30B A3B Instruct 2507$0.05$0.19OSS
20Google DeepMind logoGemma 3 12B$0.05$0.15OSS
21Google DeepMind logoGemma 3 4B$0.05$0.10OSS
22OpenAI logoGPT-5 Nano$0.05$0.40Closed
23Meta logoLlama 3.2 3B Instruct$0.05$0.33OSS
24Mistral AI logoMistral Small 3$0.05$0.08OSS
25ibm-granite logoGranite 4.2 8B$0.06$0.25Closed
26NVIDIA logoNemotron 3.5 Lightning$0.06$0.16Closed
27Alibaba Qwen logoQwen3.5-Flash$0.07$0.26OSS
28Google DeepMind logoGemma 4 26B A4B $0.07$0.23OSS
29baidu logoERNIE 4.5 21B A3B Thinking$0.07$0.28OSS
30Microsoft logoPhi 4$0.07$0.14OSS
Cheapest
Gemma 3 27B (free)
$0.00/M
$ per 1M input tokens
Why the gap

Premium reasoners pay for longer thinking budgets, better tool use, and vendor reliability. For many tasks, Gemma 3 27B (free) closes 70 to 90 percent of the GPQA gap at a fraction of the cost.

Most expensive
ERNIE 4.5 21B A3B Thinking
$0.07/M
$ per 1M input tokens
Models with explicit reasoning scores on GPQA Diamond, AIME 2024/2025, MATH-500, MMLU-Pro, HLE, DROP, BBH, or ARC-AGI. Reasoning models typically use extended chain-of-thought and burn more tokens on hard problems.