Compare · ModelsLive · 2 picked · head to head
DeepSeek V3.1 vs Kimi K2 0711
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Kimi K2 0711 wins on 3/4 benchmarks
Kimi K2 0711 wins 3 of 4 shared benchmarks. Leads in knowledge · coding.
Category leads
knowledge·Kimi K2 0711reasoning·DeepSeek V3.1coding·Kimi K2 0711
Hype vs Reality
Attention vs performance
DeepSeek V3.1
#118 by perf·no signal
Kimi K2 0711
#72 by perf·#17 by attention
Best value
DeepSeek V3.1
2.2x better value than Kimi K2 0711
DeepSeek V3.1
84.5 pts/$
$0.60/M
Kimi K2 0711
39.1 pts/$
$1.43/M
Vendor risk
Mixed exposure
One or more vendors flagged
DeepSeek
$3.4B·Tier 1
Moonshot AI
$18.0B·Tier 1
Head to head
4 benchmarks · 2 models
DeepSeek V3.1Kimi K2 0711
Fiction.LiveBench
Kimi K2 0711 leads by +8.3
Fiction.LiveBench · a continuously updated benchmark using recently published fiction to test reading comprehension and reasoning, preventing data contamination.
DeepSeek V3.1
52.8
Kimi K2 0711
61.1
Lech Mazur Writing
Kimi K2 0711 leads by +0.4
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
DeepSeek V3.1
85.2
Kimi K2 0711
85.6
SimpleBench
DeepSeek V3.1 leads by +16.4
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
DeepSeek V3.1
28.0
Kimi K2 0711
11.6
WeirdML
Kimi K2 0711 leads by +1.0
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
DeepSeek V3.1
38.4
Kimi K2 0711
39.4
Full benchmark table
| Benchmark | DeepSeek V3.1 | Kimi K2 0711 |
|---|---|---|
Fiction.LiveBench Fiction.LiveBench · a continuously updated benchmark using recently published fiction to test reading comprehension and reasoning, preventing data contamination. | 52.8 | 61.1 |
Lech Mazur Writing Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication. | 85.2 | 85.6 |
SimpleBench SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking. | 28.0 | 11.6 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 38.4 | 39.4 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.25 | $0.95 | 164K tokens (~82 books) | $4.25 | |
| $0.57 | $2.30 | 131K tokens (~66 books) | $10.03 |