Compare · ModelsLive · 2 picked · head to head
Kimi K2 Thinking vs Qwen3.6 Max Preview
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Qwen3.6 Max Preview wins on 5/6 benchmarks
Qwen3.6 Max Preview wins 5 of 6 shared benchmarks. Leads in knowledge · math.
Category leads
knowledge·Qwen3.6 Max Previewmath·Qwen3.6 Max Preview
Hype vs Reality
Attention vs performance
Kimi K2 Thinking
#103 by perf·#17 by attention
Qwen3.6 Max Preview
#112 by perf·#2 by attention
Best value
Kimi K2 Thinking
2.4x better value than Qwen3.6 Max Preview
Kimi K2 Thinking
34.0 pts/$
$1.55/M
Qwen3.6 Max Preview
14.2 pts/$
$3.59/M
Vendor risk
Who is behind the model
Moonshot AI
$18.0B·Tier 1
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
6 benchmarks · 2 models
Kimi K2 ThinkingQwen3.6 Max Preview
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Kimi K2 Thinking
15.8
Qwen3.6 Max Preview
15.8
FrontierMath-2025-02-28-Private
Qwen3.6 Max Preview leads by +3.0
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Kimi K2 Thinking
37.5
Qwen3.6 Max Preview
40.5
FrontierMath-Tier-4-2025-07-01-Private
Qwen3.6 Max Preview leads by +6.9
FrontierMath Tier 4 (Jul 2025) · the most challenging tier of frontier mathematics, containing problems that push the absolute limits of AI mathematical reasoning.
Kimi K2 Thinking
0.0
Qwen3.6 Max Preview
6.9
GPQA diamond
Qwen3.6 Max Preview leads by +4.2
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Kimi K2 Thinking
79.0
Qwen3.6 Max Preview
83.2
OTIS Mock AIME 2024-2025
Qwen3.6 Max Preview leads by +8.1
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Kimi K2 Thinking
83.0
Qwen3.6 Max Preview
91.1
SimpleQA Verified
Qwen3.6 Max Preview leads by +20.4
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Kimi K2 Thinking
31.6
Qwen3.6 Max Preview
52.0
Full benchmark table
| Benchmark | Kimi K2 Thinking | Qwen3.6 Max Preview |
|---|---|---|
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 15.8 | 15.8 |
FrontierMath-2025-02-28-Private FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning. | 37.5 | 40.5 |
FrontierMath-Tier-4-2025-07-01-Private FrontierMath Tier 4 (Jul 2025) · the most challenging tier of frontier mathematics, containing problems that push the absolute limits of AI mathematical reasoning. | 0.0 | 6.9 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 79.0 | 83.2 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 83.0 | 91.1 |
SimpleQA Verified SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information. | 31.6 | 52.0 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.60 | $2.50 | 262K tokens (~131 books) | $10.75 | |
| $1.03 | $6.16 | 262K tokens (~131 books) | $23.11 |