Licensing · Open source

Open source LLM pricing

Every open-source LLM with license, cheapest API provider, and rough self-host cost. Weights on HuggingFace, API on OpenRouter, DeepInfra, Together, Fireworks, or self-host.

Models243
Cheapest API$0.00
LicensesApache · MIT · Custom
What this page is
This page lists every open-source LLM with priced API access. For each, we show the license, cheapest API on the market, and a rough self-host cost estimate based on parameter count. Weights are downloadable from HuggingFace. APIs are offered by OpenRouter, DeepInfra, Together, Fireworks, Groq, Cerebras, and others. For very high volume, self-hosting on reserved GPUs is cheaper; for anything under 100M tokens per day, API access usually wins on total cost.

Cheapest input price first.

#ModelCheapest API
1Gemma 3 12B (free)$0.00/M
2Gemma 3 27B (free)$0.00/M
3Gemma 3 4B (free)$0.00/M
4Gemma 3n 2B (free)$0.00/M
5Gemma 3n 4B (free)$0.00/M
6Gemma 4 26B A4B (free)$0.00/M
7Gemma 4 31B (free)$0.00/M
8GLM 4.5 Air (free)$0.00/M
9gpt-oss-120b (free)$0.00/M
10gpt-oss-20b (free)$0.00/M
11Hermes 3 405B Instruct (free)$0.00/M
12Laguna M.1 (free)$0.00/M
13Laguna XS.2 (free)$0.00/M
14LFM2.5-1.2B-Instruct (free)$0.00/M
15LFM2.5-1.2B-Thinking (free)$0.00/M
16Llama 3.2 3B Instruct (free)$0.00/M
17Llama 3.3 70B Instruct (free)$0.00/M
18MiniMax M2.5 (free)$0.00/M
19Mistral Small 3.1 24B (free)$0.00/M
20Nemotron 3 Nano 30B A3B (free)$0.00/M
21Nemotron 3 Nano Omni (free)$0.00/M
22Nemotron 3 Super (free)$0.00/M
23Nemotron 3 Ultra (free)$0.00/M
24Nemotron 3.5 Content Safety (free)$0.00/M
25Nemotron Nano 12B 2 VL (free)$0.00/M
26Nemotron Nano 9B V2 (free)$0.00/M
27Nex-N2-Pro (free)$0.00/M
28North Mini Code (free)$0.00/M
29Qwen3 4B (free)$0.00/M
30Qwen3 Coder 480B A35B (free)$0.00/M
31Qwen3 Next 80B A3B Instruct (free)$0.00/M
32Qwen3.6 Plus Preview (free)$0.00/M
33Step 3.5 Flash (free)$0.00/M
34Trinity Large Preview (free)$0.00/M
35Trinity Mini (free)$0.00/M
36Uncensored (free)$0.00/M
37LFM2-2.6B$0.01/M
38LFM2-8B-A1B$0.01/M
39Granite 4.0 Micro$0.02/M
40gpt-oss-20b$0.02/M
41Mistral Nemo$0.02/M
42Llama 3.1 8B Instruct$0.02/M
43DeepSeek V4 Flash$0.03/M
44Llama 3.2 1B Instruct$0.03/M
45Gemma 2 9B$0.03/M
46LFM2-24B-A2B$0.03/M
47Qwen2.5 Coder 7B Instruct$0.03/M
48Qwen-Turbo$0.03/M
49gpt-oss-120b$0.04/M
50Llama 3 8B Lunaris$0.04/M
51Nemotron Nano 9B V2$0.04/M
52Trinity Mini$0.04/M
53Qwen3 30B A3B Instruct 2507$0.05/M
54Gemma 3 12B$0.05/M
55Gemma 3 4B$0.05/M
56Granite 4.1 8B$0.05/M
57Llama 3.2 3B Instruct$0.05/M
58Mistral Small 3$0.05/M
59Nemotron 3 Nano 30B A3B$0.05/M
60Olmo 2 32B Instruct$0.05/M
61Gemma 3n 4B$0.06/M
62GLM 4.7 Flash$0.06/M
63Hy3 preview$0.06/M
64Qwen3.5-Flash$0.07/M
65Gemma 4 26B A4B $0.07/M
66ERNIE 4.5 21B A3B$0.07/M
67ERNIE 4.5 21B A3B Thinking$0.07/M
68Phi 4$0.07/M
69Qwen3 Coder 30B A3B Instruct$0.07/M
70gpt-oss-safeguard-20b$0.07/M
71Gemma 3 27B$0.08/M
72MythoMax 13B$0.08/M
73Nemotron 3 Super$0.08/M
74Phi 4 Mini Instruct$0.08/M
75Qwen3 32B$0.08/M
76Gemma 4 31B$0.09/M
77MiMo-V2-Flash$0.09/M
78Qwen3 235B A22B Instruct 2507$0.09/M
79Qwen3 Next 80B A3B Instruct$0.09/M
80Tongyi DeepResearch 30B A3B$0.09/M
81Mistral Small 3.2 24B$0.09/M
82Devstral Small 1.1$0.10/M
83Laguna XS.2$0.10/M
84Llama 3.3 70B Instruct$0.10/M
85Llama 4 Scout$0.10/M
86Ministral 3 3B 2512$0.10/M
87Mistral Small Creative$0.10/M
88Qwen2.5 7B Instruct$0.10/M
89Qwen3.5-9B$0.10/M
90Reka Edge$0.10/M
91Reka Flash 3$0.10/M
92Step 3.5 Flash$0.10/M
93UI-TARS 7B $0.10/M
94Voxtral Small 24B 2507$0.10/M
95Qwen3 VL 32B Instruct$0.10/M
96Mistral 7B Instruct v0.1$0.11/M
97Qwen3 8B$0.12/M
98Qwen3 VL 8B Instruct$0.12/M
99Qwen3 14B$0.12/M
100Qwen3 30B A3B$0.12/M
101Qwen3 Coder Next$0.12/M
102GLM 4.5 Air$0.13/M
103Hermes 4 70B$0.13/M
104DeepSeek V3.1 Nex N1$0.14/M
105Qwen VL Plus$0.14/M
106ERNIE 4.5 VL 28B A3B$0.14/M
107Hermes 2 Pro - Llama-3 8B$0.14/M
108Hunyuan A13B Instruct$0.14/M
109Llama 3 8B Instruct$0.14/M
110MiMo-V2.5$0.14/M
111Ministral 3 8B 2512$0.15/M
112Mistral Small 4$0.15/M
113Olmo 3 32B Think$0.15/M
114Olmo 3.1 32B Think$0.15/M
115Qwen3 Next 80B A3B Thinking$0.15/M
116Qwen3 VL 30B A3B Instruct$0.15/M
117Qwen3.5-35B-A3B$0.15/M
118Qwen3.6 35B A3B$0.15/M
119QwQ 32B$0.15/M
120Rnj 1 Instruct$0.15/M
121Llama Guard 4 12B$0.18/M
122Qwen3 VL 8B Thinking$0.18/M
123Llama 4 Maverick$0.19/M
124Qwen3.6 Flash$0.19/M
125Qwen3 Coder Flash$0.20/M
126Qwen3.5-27B$0.20/M
127INTELLECT-3$0.20/M
128Laguna M.1$0.20/M
129LongCat Flash Chat$0.20/M
130MiniMax-01$0.20/M
131Ministral 3 14B 2512$0.20/M
132Nemotron Nano 12B 2 VL$0.20/M
133Olmo 3.1 32B Instruct$0.20/M
134Qwen2.5 VL 32B Instruct$0.20/M
135Qwen3 30B A3B Thinking 2507$0.20/M
136Qwen3 VL 30B A3B Thinking$0.20/M
137Saba$0.20/M
138Step 3.7 Flash$0.20/M
139DeepSeek V4 Pro$0.21/M
140MiniMax M2.7$0.21/M
141Qwen3 VL 235B A22B Instruct$0.21/M
142Qwen3 235B A22B Thinking 2507$0.23/M
143DeepSeek V3 0324$0.25/M
144DeepSeek V3.1$0.25/M
145Rocinante 12B$0.25/M
146Trinity Large Thinking$0.25/M
147DeepSeek V3$0.26/M
148Qwen Plus 0728$0.26/M
149Qwen Plus 0728 (thinking)$0.26/M
150Qwen-Plus$0.26/M
151Qwen3.5 Plus 2026-02-15$0.26/M
152Qwen3.5-122B-A10B$0.26/M
153DeepSeek V3.1 Terminus$0.27/M
154DeepSeek V3.2 Exp$0.27/M
155MiniMax M2.5$0.27/M
156DeepSeek V3.2$0.28/M
157ERNIE 4.5 300B A47B $0.28/M
158R1 Distill Qwen 32B$0.29/M
159Codestral 2508$0.30/M
160Cydonia 24B V4.1$0.30/M
161DeepSeek R1T2 Chimera$0.30/M
162GLM 4.6V$0.30/M
163MiniMax M2$0.30/M
164MiniMax M2.1$0.30/M
165MiniMax M3$0.30/M
166Qwen3 Coder 480B A35B$0.30/M
167Qwen3.5 Plus 2026-04-20$0.30/M
168Qwen3.6 27B$0.32/M
169Qwen3.7 Plus$0.32/M
170Qwen3.6 Plus$0.33/M
171Llama 3.2 11B Vision Instruct$0.34/M
172ReMM SLERP 13B$0.35/M
173Mistral Small 3.1 24B$0.35/M
174Qwen2.5 72B Instruct$0.36/M
175DeepSeek V3.2 Speciale$0.40/M
176Devstral 2 2512$0.40/M
177Devstral Medium$0.40/M
178Llama 3.1 70B Instruct$0.40/M
179Llama 3.3 Nemotron Super 49B V1.5$0.40/M
180Mistral Medium 3$0.40/M
181Mistral Medium 3.1$0.40/M
182Qwen3 VL 235B A22B Thinking$0.40/M
183UnslopNemo 12B$0.40/M
184ERNIE 4.5 VL 424B A47B $0.42/M
185GLM 4.6$0.43/M
186MiMo-V2.5-Pro$0.43/M
187Kimi K2.5$0.45/M
188Qwen3 235B A22B$0.46/M
189Llama Guard 3 8B$0.48/M
190Mistral Large 3 2512$0.50/M
191Nemotron 3 Ultra$0.50/M
192Nex-N2-Pro$0.50/M
193R1 0528$0.50/M
194Llama 3 70B Instruct$0.51/M
195Qwen VL Max$0.52/M
196Mixtral 8x7B Instruct$0.54/M
197Qwen3.5 397B A17B$0.55/M
198Skyfall 36B V2$0.55/M
199Kimi K2 0711$0.57/M
200GLM 4.5$0.60/M
201GLM 4.5V$0.60/M
202GLM 4.7$0.60/M
203GLM 5$0.60/M
204Kimi K2 0905$0.60/M
205Kimi K2 Thinking$0.60/M
206Llama 3.1 Nemotron Ultra 253B v1$0.60/M
207Kimi K2.7 Code$0.61/M
208WizardLM-2 8x22B$0.62/M
209Gemma 2 27B$0.65/M
210Llama 3.3 Euryale 70B$0.65/M
211Qwen3 Coder Plus$0.65/M
212Qwen2.5 Coder 32B Instruct$0.66/M
213Aion-1.0-Mini$0.70/M
214Hermes 3 70B Instruct$0.70/M
215R1$0.70/M
216Qwen3 Max$0.78/M
217Qwen3 Max Thinking$0.78/M
218CodeLLaMa 7B Instruct Solidity$0.80/M
219Llemma 7b$0.80/M
220Qwen2.5 VL 72B Instruct$0.80/M
221R1 Distill Llama 70B$0.80/M
222Llama 3.1 Euryale 70B v2.2$0.85/M
223Kimi K2.6$0.95/M
224GLM 5.1$0.97/M
225GLM 5.2$1.00/M
226Hermes 3 405B Instruct$1.00/M
227Hermes 4 405B$1.00/M
228Qwen3.6 Max Preview$1.03/M
229Qwen-Max $1.04/M
230Llama 3.1 Nemotron 70B Instruct$1.20/M
231Qwen3.7 Max$1.25/M
232Llama 3 Euryale 70B v2.1$1.48/M
233Mistral Medium 3.5$1.50/M
234Jamba Large 1.7$2.00/M
235Mistral Large$2.00/M
236Mistral Large 2407$2.00/M
237Mistral Large 2411$2.00/M
238Mixtral 8x22B Instruct$2.00/M
239Pixtral Large 2411$2.00/M
240Command A$2.50/M
241Llama 3.1 70B Hanami x1$3.00/M
242Magnum v4 72B$3.00/M
243Goliath 120B$3.75/M
Low volume
Under 1M tokens/day

Always pick API. A single GPU hour wipes out weeks of API spend at this volume.

Mid volume
10M to 100M tokens/day

Depends on model size. Small models (under 30B) are cheaper via API. Large models (200B+) favor dedicated GPUs.

High volume
1B+ tokens/day

Self-host wins, if utilization stays near 100%. Use reserved instances and bundle across workloads.

Cheapest
Gemma 3 12B (free)
$0.00/M
$ per 1M input tokens
Why the gap

Within OSS, the price range is driven by parameter count (bigger = more expensive to serve) and provider margins. Smaller dense models and heavily quantized serves sit at the bottom.

Most expensive
Goliath 120B
$3.75/M
$ per 1M input tokens
The weights are publicly downloadable under some license (Apache 2.0, MIT, Llama Community, Qwen License, DeepSeek License, etc.). Not every "open" license is actually OSI-compliant · always read the license before commercial use.