Context · 32K+

Cheapest 32K context LLMs

Every LLM with at least 32K token context. Ranked by input price per 1M tokens.

Models50
Cheapest$0.00
Min context32K tokens
What this page is
32K context is the floor for modern LLM use. This tier captures the broadest set of priced models at one of the deepest discounts. Ideal for chat, short-doc RAG, classification, and any workload where you do not need a huge window.

32K+ context models, cheapest first.

#ModelIn $/1MOut $/1MType
1DeepSeek logoDeepSeek V4 Flash 0731 (free)$0.00$0.00Closed
2openrouter logoElephant$0.00$0.00Closed
3openrouter logoFree Models Router$0.00$0.00Closed
4Google DeepMind logoGemma 3 12B (free)$0.00$0.00OSS
5Google DeepMind logoGemma 3 27B (free)$0.00$0.00OSS
6Google DeepMind logoGemma 3 4B (free)$0.00$0.00OSS
7Google DeepMind logoGemma 4 26B A4B (free)$0.00$0.00OSS
8Google DeepMind logoGemma 4 31B (free)$0.00$0.00OSS
9z-ai logoGLM 4.5 Air (free)$0.00$0.00OSS
10z-ai logoGLM 5.2 (free)$0.00$0.00Closed
11OpenAI logogpt-oss-120b (free)$0.00$0.00OSS
12OpenAI logogpt-oss-20b (free)$0.00$0.00OSS
13nousresearch logoHermes 3 405B Instruct (free)$0.00$0.00OSS
14tencent logoHy3 (free)$0.00$0.00Closed
15tencent logoHy3 preview (free)$0.00$0.00Closed
16Laguna M.1 (free)$0.00$0.00OSS
17Laguna XS.2 (free)$0.00$0.00OSS
18liquid logoLFM2.5-1.2B-Instruct (free)$0.00$0.00OSS
19liquid logoLFM2.5-1.2B-Thinking (free)$0.00$0.00OSS
20liquid logoLFM2.5-2.6B (free)$0.00$0.00Closed
21Meta logoLlama 3.2 3B Instruct (free)$0.00$0.00OSS
22Meta logoLlama 3.3 70B Instruct (free)$0.00$0.00OSS
23Meta logoLlama Guard 4 12B (free)$0.00$0.00Closed
24Google DeepMind logoLyria 3 Clip Preview$0.00$0.00Closed
25Google DeepMind logoLyria 3 Pro Preview$0.00$0.00Closed
26minimax logoMiniMax M2.5 (free)$0.00$0.00OSS
27minimax logoMiniMax M2.7 (free)$0.00$0.00Closed
28minimax logoMiniMax M3 (free)$0.00$0.00Closed
29Mistral AI logoMistral Small 3.1 24B (free)$0.00$0.00OSS
30NVIDIA logoNemotron 3 Nano 30B A3B (free)$0.00$0.00OSS
31NVIDIA logoNemotron 3 Nano Omni (free)$0.00$0.00OSS
32NVIDIA logoNemotron 3 Super (free)$0.00$0.00OSS
33NVIDIA logoNemotron 3 Ultra (free)$0.00$0.00OSS
34NVIDIA logoNemotron 3.5 Content Safety (free)$0.00$0.00OSS
35NVIDIA logoNemotron 3.5 Lightning (free)$0.00$0.00Closed
36NVIDIA logoNemotron Nano 12B 2 VL (free)$0.00$0.00OSS
37NVIDIA logoNemotron Nano 9B V2 (free)$0.00$0.00OSS
38nex-agi logoNex-N2-Pro (free)$0.00$0.00OSS
39nex-agi logoNex-N2.5-Mini (free)$0.00$0.00Closed
40nex-agi logoNex-N2.5-Pro (free)$0.00$0.00Closed
41Cohere logoNorth Mini Code (free)$0.00$0.00OSS
42openrouter logoOwl Alpha$0.00$0.00Closed
43baidu logoQianfan-OCR-Fast (free)$0.00$0.00Closed
44Alibaba Qwen logoQwen3 4B (free)$0.00$0.00OSS
45Alibaba Qwen logoQwen3 Coder 480B A35B (free)$0.00$0.00OSS
46Alibaba Qwen logoQwen3 Next 80B A3B Instruct (free)$0.00$0.00OSS
47Alibaba Qwen logoQwen3.6 Plus (free)$0.00$0.00Closed
48Alibaba Qwen logoQwen3.6 Plus Preview (free)$0.00$0.00OSS
49Alibaba Qwen logoQwen3.8 27B (free)$0.00$0.00Closed
50stepfun logoStep 3.5 Flash (free)$0.00$0.00OSS
Cheapest
DeepSeek V4 Flash 0731 (free)
$0.00/M
$ per 1M input tokens
Why the gap

At this tier, price reflects raw model quality more than context size. The cheapest 32K model is often a small open-source model; the most expensive is usually a frontier model deliberately billed the same across windows.

Most expensive
DeepSeek V4 Flash 0731 (free)
$0.00/M
$ per 1M input tokens
For most applications, yes. Chat histories, RAG retrievals, and single-document QA rarely exceed 16K tokens. 32K leaves plenty of headroom.