Compare · ModelsLive · 2 picked · head to head

Llama 3.3 70B Instruct vs Qwen2.5 32B Instruct

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Qwen2.5 32B Instruct wins 5 of 9 shared benchmarks. Leads in knowledge · math.

Category leads
knowledge·Qwen2.5 32B Instructgeneral·Llama 3.3 70B Instructlanguage·Llama 3.3 70B Instructmath·Qwen2.5 32B Instructreasoning·Llama 3.3 70B Instruct
Hype vs Reality
Llama 3.3 70B Instruct
#218 by perf·#18 by attention
QUIET
Qwen2.5 32B Instruct
#222 by perf·#2 by attention
OVERHYPED
Best value
Llama 3.3 70B Instruct
171.9 pts/$
$0.21/M
Qwen2.5 32B Instruct
n/a
no price
Vendor risk
Meta logo
Meta AI
$1.87T·Tier 1
Low risk
Alibaba logo
Alibaba (Qwen)
$293.0B·Tier 1
Low risk
Head to head
Llama 3.3 70B InstructQwen2.5 32B Instruct
GPQA diamond
Llama 3.3 70B Instruct leads by +1.8
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Llama 3.3 70B Instruct
29.9
Qwen2.5 32B Instruct
28.1
BBH (HuggingFace)
Llama 3.3 70B Instruct leads by +0.1
Llama 3.3 70B Instruct
56.6
Qwen2.5 32B Instruct
56.5
GPQA
Qwen2.5 32B Instruct leads by +1.2
Llama 3.3 70B Instruct
10.5
Qwen2.5 32B Instruct
11.7
IFEval
Llama 3.3 70B Instruct leads by +6.5
Llama 3.3 70B Instruct
90.0
Qwen2.5 32B Instruct
83.5
MATH Level 5
Qwen2.5 32B Instruct leads by +14.2
Llama 3.3 70B Instruct
48.3
Qwen2.5 32B Instruct
62.5
MMLU-PRO
Qwen2.5 32B Instruct leads by +3.7
Llama 3.3 70B Instruct
48.1
Qwen2.5 32B Instruct
51.9
MUSR
Llama 3.3 70B Instruct leads by +2.1
Llama 3.3 70B Instruct
15.6
Qwen2.5 32B Instruct
13.5
MATH level 5
Qwen2.5 32B Instruct leads by +14.5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Llama 3.3 70B Instruct
41.6
Qwen2.5 32B Instruct
56.1
OTIS Mock AIME 2024-2025
Qwen2.5 32B Instruct leads by +2.2
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Llama 3.3 70B Instruct
5.0
Qwen2.5 32B Instruct
7.3
Full benchmark table
BenchmarkLlama 3.3 70B InstructQwen2.5 32B Instruct
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
29.928.1
BBH (HuggingFace)
56.656.5
GPQA
10.511.7
IFEval
90.083.5
MATH Level 5
48.362.5
MMLU-PRO
48.151.9
MUSR
15.613.5
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
41.656.1
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
5.07.3
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Meta logoLlama 3.3 70B Instruct$0.10$0.32131K tokens (~66 books)$1.55
Alibaba logoQwen2.5 32B Instruct————