Context · 32K+
Cheapest 32K context LLMs
Every LLM with at least 32K token context. Ranked by input price per 1M tokens.
Models50
Cheapest$0.00
Min context32K tokens
What this page is
32K context is the floor for modern LLM use. This tier captures the broadest set of priced models at one of the deepest discounts. Ideal for chat, short-doc RAG, classification, and any workload where you do not need a huge window.
Ranked by input price
32K+ context models, cheapest first.
Top 3 cheapest 32K+ context LLMs
Cheapest 32K
DeepSeek V4 Flash 0731 (free)
input
$0.00/M
output
$0.00/M
DeepSeek V4 Flash 0731 (free) at $0.00/M input with 1.0M context · fine for most chat and RAG loads.
Runner up
Elephant
input
$0.00/M
output
$0.00/M
Elephant at $0.00/M input with 262K context · fine for most chat and RAG loads.
Third
Free Models Router
input
$0.00/M
output
$0.00/M
Free Models Router at $0.00/M input with 200K context · fine for most chat and RAG loads.
The price gap · cheapest vs most expensive
Cheapest
DeepSeek V4 Flash 0731 (free)
$0.00/M
$ per 1M input tokens
Why the gap
At this tier, price reflects raw model quality more than context size. The cheapest 32K model is often a small open-source model; the most expensive is usually a frontier model deliberately billed the same across windows.
Most expensive
DeepSeek V4 Flash 0731 (free)
$0.00/M
$ per 1M input tokens
Frequently asked questions
For most applications, yes. Chat histories, RAG retrievals, and single-document QA rarely exceed 16K tokens. 32K leaves plenty of headroom.