Compare · ModelsLive · 2 picked · head to head

GPT-4o-mini (2024-07-18) vs QwQ 32B

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

QwQ 32B wins 5 of 5 shared benchmarks. Leads in coding · arena · knowledge.

Category leads
coding·QwQ 32Barena·QwQ 32Bknowledge·QwQ 32Bmath·QwQ 32B
Hype vs Reality
GPT-4o-mini (2024-07-18)
#171 by perf·no signal
QUIET
QwQ 32B
#247 by perf·no signal
QUIET
Best value
1.4x better value than QwQ 32B
GPT-4o-mini (2024-07-18)
115.2 pts/$
$0.38/M
QwQ 32B
84.7 pts/$
$0.36/M
Vendor risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Alibaba Qwen logo
Alibaba (Qwen)
$293.0B·Tier 1
Low risk
Head to head
GPT-4o-mini (2024-07-18)QwQ 32B
Aider polyglot
QwQ 32B leads by +17.3
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
GPT-4o-mini (2024-07-18)
3.6
QwQ 32B
20.9
Chatbot Arena Elo · Overall
QwQ 32B leads by +18.1
GPT-4o-mini (2024-07-18)
1317.7
QwQ 32B
1335.8
GPQA diamond
QwQ 32B leads by +36.8
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
GPT-4o-mini (2024-07-18)
17.0
QwQ 32B
53.8
Lech Mazur Writing
QwQ 32B leads by +13.0
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
GPT-4o-mini (2024-07-18)
67.2
QwQ 32B
80.2
OTIS Mock AIME 2024-2025
QwQ 32B leads by +52.3
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
GPT-4o-mini (2024-07-18)
6.8
QwQ 32B
59.1
Full benchmark table
BenchmarkGPT-4o-mini (2024-07-18)QwQ 32B
Aider polyglot
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
3.620.9
Chatbot Arena Elo · Overall
1317.71335.8
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
17.053.8
Lech Mazur Writing
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
67.280.2
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
6.859.1
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
OpenAI logoGPT-4o-mini (2024-07-18)$0.15$0.60128K tokens (~64 books)$2.62
Alibaba Qwen logoQwQ 32B$0.15$0.58131K tokens (~66 books)$2.57