Context · 128K+

Cheapest 128K context LLMs

Every LLM with a 128,000+ token context window. Ranked by input price per 1M tokens.

Models40
Cheapest$0.00
Min context128K tokens
What this page is
128K is the modern baseline for LLM context. Every model priced per 1M tokens at 128K+ is listed here, cheapest first. For RAG pipelines, long conversations, and mid-sized documents, this is the sweet spot.

128K+ context models, cheapest first.

#ModelIn $/1MOut $/1MType
1DeepSeek logoDeepSeek V4 Flash 0731 (free)$0.00$0.00Closed
2openrouter logoElephant$0.00$0.00Closed
3openrouter logoFree Models Router$0.00$0.00Closed
4Google DeepMind logoGemma 3 27B (free)$0.00$0.00OSS
5Google DeepMind logoGemma 4 26B A4B (free)$0.00$0.00OSS
6Google DeepMind logoGemma 4 31B (free)$0.00$0.00OSS
7z-ai logoGLM 4.5 Air (free)$0.00$0.00OSS
8OpenAI logogpt-oss-120b (free)$0.00$0.00OSS
9OpenAI logogpt-oss-20b (free)$0.00$0.00OSS
10nousresearch logoHermes 3 405B Instruct (free)$0.00$0.00OSS
11tencent logoHy3 (free)$0.00$0.00Closed
12tencent logoHy3 preview (free)$0.00$0.00Closed
13Laguna M.1 (free)$0.00$0.00OSS
14Laguna XS.2 (free)$0.00$0.00OSS
15Meta logoLlama 3.2 3B Instruct (free)$0.00$0.00OSS
16Meta logoLlama 3.3 70B Instruct (free)$0.00$0.00OSS
17Meta logoLlama Guard 4 12B (free)$0.00$0.00Closed
18Google DeepMind logoLyria 3 Clip Preview$0.00$0.00Closed
19Google DeepMind logoLyria 3 Pro Preview$0.00$0.00Closed
20minimax logoMiniMax M2.5 (free)$0.00$0.00OSS
21minimax logoMiniMax M2.7 (free)$0.00$0.00Closed
22minimax logoMiniMax M3 (free)$0.00$0.00Closed
23Mistral AI logoMistral Small 3.1 24B (free)$0.00$0.00OSS
24NVIDIA logoNemotron 3 Nano 30B A3B (free)$0.00$0.00OSS
25NVIDIA logoNemotron 3 Nano Omni (free)$0.00$0.00OSS
26NVIDIA logoNemotron 3 Super (free)$0.00$0.00OSS
27NVIDIA logoNemotron 3 Ultra (free)$0.00$0.00OSS
28NVIDIA logoNemotron 3.5 Content Safety (free)$0.00$0.00OSS
29NVIDIA logoNemotron 3.5 Lightning (free)$0.00$0.00Closed
30NVIDIA logoNemotron Nano 12B 2 VL (free)$0.00$0.00OSS
31NVIDIA logoNemotron Nano 9B V2 (free)$0.00$0.00OSS
32nex-agi logoNex-N2-Pro (free)$0.00$0.00OSS
33nex-agi logoNex-N2.5-Mini (free)$0.00$0.00Closed
34nex-agi logoNex-N2.5-Pro (free)$0.00$0.00Closed
35Cohere logoNorth Mini Code (free)$0.00$0.00OSS
36openrouter logoOwl Alpha$0.00$0.00Closed
37Alibaba Qwen logoQwen3 Coder 480B A35B (free)$0.00$0.00OSS
38Alibaba Qwen logoQwen3 Next 80B A3B Instruct (free)$0.00$0.00OSS
39Alibaba Qwen logoQwen3.6 Plus (free)$0.00$0.00Closed
40Alibaba Qwen logoQwen3.6 Plus Preview (free)$0.00$0.00OSS
Cheapest
DeepSeek V4 Flash 0731 (free)
$0.00/M
$ per 1M input tokens
Why the gap

At 128K, premium pricing pays for reasoning quality and vendor reliability, not window size. For RAG backbones, the cheap end almost always wins.

Most expensive
DeepSeek V4 Flash 0731 (free)
$0.00/M
$ per 1M input tokens
For most production use cases, yes. 128K fits 300+ pages, a full API spec, or a large function library. Only reach for 200K+ when you regularly exceed 100K token prompts.