Compare · ModelsLive · 2 picked · head to head
DeepSeek R1 Distill Qwen 14B vs Qwen2.5 72B Instruct
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Qwen2.5 72B Instruct wins on 5/9 benchmarks
Qwen2.5 72B Instruct wins 5 of 9 shared benchmarks. Leads in knowledge · general · language.
Category leads
knowledge·Qwen2.5 72B Instructgeneral·Qwen2.5 72B Instructlanguage·Qwen2.5 72B Instructmath·DeepSeek R1 Distill Qwen 14Breasoning·DeepSeek R1 Distill Qwen 14B
Hype vs Reality
Attention vs performance
DeepSeek R1 Distill Qwen 14B
#107 by perf·no signal
Qwen2.5 72B Instruct
#128 by perf·#2 by attention
Best value
Qwen2.5 72B Instruct
DeepSeek R1 Distill Qwen 14B
n/a
no price
Qwen2.5 72B Instruct
129.5 pts/$
$0.38/M
Vendor risk
Mixed exposure
One or more vendors flagged
DeepSeek
$3.4B·Tier 1
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
9 benchmarks · 2 models
DeepSeek R1 Distill Qwen 14BQwen2.5 72B Instruct
GPQA diamond
Qwen2.5 72B Instruct leads by +5.9
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
DeepSeek R1 Distill Qwen 14B
26.3
Qwen2.5 72B Instruct
32.2
BBH (HuggingFace)
Qwen2.5 72B Instruct leads by +21.2
DeepSeek R1 Distill Qwen 14B
40.7
Qwen2.5 72B Instruct
61.9
GPQA
DeepSeek R1 Distill Qwen 14B leads by +1.7
DeepSeek R1 Distill Qwen 14B
18.3
Qwen2.5 72B Instruct
16.7
IFEval
Qwen2.5 72B Instruct leads by +42.6
DeepSeek R1 Distill Qwen 14B
43.8
Qwen2.5 72B Instruct
86.4
MATH Level 5
Qwen2.5 72B Instruct leads by +2.8
DeepSeek R1 Distill Qwen 14B
57.0
Qwen2.5 72B Instruct
59.8
MMLU-PRO
Qwen2.5 72B Instruct leads by +10.7
DeepSeek R1 Distill Qwen 14B
40.7
Qwen2.5 72B Instruct
51.4
MUSR
DeepSeek R1 Distill Qwen 14B leads by +17.0
DeepSeek R1 Distill Qwen 14B
28.7
Qwen2.5 72B Instruct
11.7
MATH level 5
DeepSeek R1 Distill Qwen 14B leads by +24.0
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
DeepSeek R1 Distill Qwen 14B
87.1
Qwen2.5 72B Instruct
63.2
OTIS Mock AIME 2024-2025
DeepSeek R1 Distill Qwen 14B leads by +42.5
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
DeepSeek R1 Distill Qwen 14B
50.5
Qwen2.5 72B Instruct
8.0
Full benchmark table
| Benchmark | DeepSeek R1 Distill Qwen 14B | Qwen2.5 72B Instruct |
|---|---|---|
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 26.3 | 32.2 |
BBH (HuggingFace) | 40.7 | 61.9 |
GPQA | 18.3 | 16.7 |
IFEval | 43.8 | 86.4 |
MATH Level 5 | 57.0 | 59.8 |
MMLU-PRO | 40.7 | 51.4 |
MUSR | 28.7 | 11.7 |
MATH level 5 MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics. | 87.1 | 63.2 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 50.5 | 8.0 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| — | — | — | — | |
| $0.36 | $0.40 | 33K tokens (~16 books) | $3.70 |