Compare · ModelsLive · 2 picked · head to head

Kimi K2 Thinking vs Qwen3.6 Max Preview

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Qwen3.6 Max Preview wins 5 of 6 shared benchmarks. Leads in knowledge · math.

Category leads
knowledge·Qwen3.6 Max Previewmath·Qwen3.6 Max Preview
Hype vs Reality
Kimi K2 Thinking
#103 by perf·#17 by attention
UNDERRATED
Qwen3.6 Max Preview
#112 by perf·#2 by attention
DESERVED
Best value
2.4x better value than Qwen3.6 Max Preview
Kimi K2 Thinking
34.0 pts/$
$1.55/M
Qwen3.6 Max Preview
14.2 pts/$
$3.59/M
Vendor risk
moonshotai logo
Moonshot AI
$18.0B·Tier 1
Medium risk
Alibaba Qwen logo
Alibaba (Qwen)
$293.0B·Tier 1
Low risk
Head to head
Kimi K2 ThinkingQwen3.6 Max Preview
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Kimi K2 Thinking
15.8
Qwen3.6 Max Preview
15.8
FrontierMath-2025-02-28-Private
Qwen3.6 Max Preview leads by +3.0
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Kimi K2 Thinking
37.5
Qwen3.6 Max Preview
40.5
FrontierMath-Tier-4-2025-07-01-Private
Qwen3.6 Max Preview leads by +6.9
FrontierMath Tier 4 (Jul 2025) · the most challenging tier of frontier mathematics, containing problems that push the absolute limits of AI mathematical reasoning.
Kimi K2 Thinking
0.0
Qwen3.6 Max Preview
6.9
GPQA diamond
Qwen3.6 Max Preview leads by +4.2
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Kimi K2 Thinking
79.0
Qwen3.6 Max Preview
83.2
OTIS Mock AIME 2024-2025
Qwen3.6 Max Preview leads by +8.1
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Kimi K2 Thinking
83.0
Qwen3.6 Max Preview
91.1
SimpleQA Verified
Qwen3.6 Max Preview leads by +20.4
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Kimi K2 Thinking
31.6
Qwen3.6 Max Preview
52.0
Full benchmark table
BenchmarkKimi K2 ThinkingQwen3.6 Max Preview
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
15.815.8
FrontierMath-2025-02-28-Private
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
37.540.5
FrontierMath-Tier-4-2025-07-01-Private
FrontierMath Tier 4 (Jul 2025) · the most challenging tier of frontier mathematics, containing problems that push the absolute limits of AI mathematical reasoning.
0.06.9
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
79.083.2
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
83.091.1
SimpleQA Verified
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
31.652.0
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
moonshotai logoKimi K2 Thinking$0.60$2.50262K tokens (~131 books)$10.75
Alibaba Qwen logoQwen3.6 Max Preview$1.03$6.16262K tokens (~131 books)$23.11