Compare · ModelsLive · 2 picked · head to head
Llama 3.3 70B Instruct vs Qwen2.5 32B Instruct
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Qwen2.5 32B Instruct wins on 5/9 benchmarks
Qwen2.5 32B Instruct wins 5 of 9 shared benchmarks. Leads in knowledge · math.
Category leads
knowledge·Qwen2.5 32B Instructgeneral·Llama 3.3 70B Instructlanguage·Llama 3.3 70B Instructmath·Qwen2.5 32B Instructreasoning·Llama 3.3 70B Instruct
Hype vs Reality
Attention vs performance
Llama 3.3 70B Instruct
#218 by perf·#18 by attention
Qwen2.5 32B Instruct
#222 by perf·#2 by attention
Best value
Llama 3.3 70B Instruct
Llama 3.3 70B Instruct
171.9 pts/$
$0.21/M
Qwen2.5 32B Instruct
n/a
no price
Vendor risk
Who is behind the model
Meta AI
$1.87T·Tier 1
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
9 benchmarks · 2 models
Llama 3.3 70B InstructQwen2.5 32B Instruct
GPQA diamond
Llama 3.3 70B Instruct leads by +1.8
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Llama 3.3 70B Instruct
29.9
Qwen2.5 32B Instruct
28.1
BBH (HuggingFace)
Llama 3.3 70B Instruct leads by +0.1
Llama 3.3 70B Instruct
56.6
Qwen2.5 32B Instruct
56.5
GPQA
Qwen2.5 32B Instruct leads by +1.2
Llama 3.3 70B Instruct
10.5
Qwen2.5 32B Instruct
11.7
IFEval
Llama 3.3 70B Instruct leads by +6.5
Llama 3.3 70B Instruct
90.0
Qwen2.5 32B Instruct
83.5
MATH Level 5
Qwen2.5 32B Instruct leads by +14.2
Llama 3.3 70B Instruct
48.3
Qwen2.5 32B Instruct
62.5
MMLU-PRO
Qwen2.5 32B Instruct leads by +3.7
Llama 3.3 70B Instruct
48.1
Qwen2.5 32B Instruct
51.9
MUSR
Llama 3.3 70B Instruct leads by +2.1
Llama 3.3 70B Instruct
15.6
Qwen2.5 32B Instruct
13.5
MATH level 5
Qwen2.5 32B Instruct leads by +14.5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Llama 3.3 70B Instruct
41.6
Qwen2.5 32B Instruct
56.1
OTIS Mock AIME 2024-2025
Qwen2.5 32B Instruct leads by +2.2
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Llama 3.3 70B Instruct
5.0
Qwen2.5 32B Instruct
7.3
Full benchmark table
| Benchmark | Llama 3.3 70B Instruct | Qwen2.5 32B Instruct |
|---|---|---|
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 29.9 | 28.1 |
BBH (HuggingFace) | 56.6 | 56.5 |
GPQA | 10.5 | 11.7 |
IFEval | 90.0 | 83.5 |
MATH Level 5 | 48.3 | 62.5 |
MMLU-PRO | 48.1 | 51.9 |
MUSR | 15.6 | 13.5 |
MATH level 5 MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics. | 41.6 | 56.1 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 5.0 | 7.3 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.10 | $0.32 | 131K tokens (~66 books) | $1.55 | |
| — | — | — | — |