Compare · ModelsLive · 2 picked · head to head
Grok 4.5 vs Qwen3.6 Plus
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Grok 4.5 wins on 9/9 benchmarks
Grok 4.5 wins 9 of 9 shared benchmarks. Leads in arena · knowledge · general.
Category leads
arena·Grok 4.5knowledge·Grok 4.5general·Grok 4.5math·Grok 4.5
Hype vs Reality
Attention vs performance
Grok 4.5
#88 by perf·#15 by attention
Qwen3.6 Plus
#94 by perf·#2 by attention
Best value
Qwen3.6 Plus
3.5x better value than Grok 4.5
Grok 4.5
13.6 pts/$
$4.00/M
Qwen3.6 Plus
47.0 pts/$
$1.14/M
Vendor risk
Who is behind the model
xAI
$250.0B·Tier 1
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
9 benchmarks · 2 models
Grok 4.5Qwen3.6 Plus
Chatbot Arena Elo · Coding
Grok 4.5 leads by +91.1
Grok 4.5
1551.8
Qwen3.6 Plus
1460.7
Chatbot Arena Elo · Overall
Grok 4.5 leads by +22.4
Grok 4.5
1465.8
Qwen3.6 Plus
1443.4
Chess Puzzles
Grok 4.5 leads by +20.0
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Grok 4.5
32.7
Qwen3.6 Plus
12.7
Dtbench
Grok 4.5 leads by +24.4
Grok 4.5
94.2
Qwen3.6 Plus
69.8
FrontierMath-Tiers-1-3-v2-Private
Grok 4.5 leads by +18.9
Grok 4.5
57.2
Qwen3.6 Plus
38.3
GPQA diamond
Grok 4.5 leads by +6.7
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Grok 4.5
91.3
Qwen3.6 Plus
84.5
Lmca
Grok 4.5 leads by +14.3
Grok 4.5
53.2
Qwen3.6 Plus
38.9
OTIS Mock AIME 2024-2025
Grok 4.5 leads by +4.5
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Grok 4.5
97.8
Qwen3.6 Plus
93.3
SimpleQA Verified
Grok 4.5 leads by +4.2
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Grok 4.5
48.3
Qwen3.6 Plus
44.1
Full benchmark table
| Benchmark | Grok 4.5 | Qwen3.6 Plus |
|---|---|---|
Chatbot Arena Elo · Coding | 1551.8 | 1460.7 |
Chatbot Arena Elo · Overall | 1465.8 | 1443.4 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 32.7 | 12.7 |
Dtbench | 94.2 | 69.8 |
FrontierMath-Tiers-1-3-v2-Private | 57.2 | 38.3 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 91.3 | 84.5 |
Lmca | 53.2 | 38.9 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 97.8 | 93.3 |
SimpleQA Verified SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information. | 48.3 | 44.1 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $2.00 | $6.00 | 500K tokens (~250 books) | $30.00 | |
| $0.33 | $1.95 | 1.0M tokens (~500 books) | $7.31 |