Compare · ModelsLive · 3 picked · head to head

Llama 3.1 70B Instruct vs Llama 3.3 70B Instruct vs Qwen2.5 72B Instruct

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Qwen2.5 72B Instruct wins 11 of 18 shared benchmarks. Leads in coding · knowledge · general.

Category leads
coding·Qwen2.5 72B Instructarena·Llama 3.3 70B Instructknowledge·Qwen2.5 72B Instructgeneral·Qwen2.5 72B Instructlanguage·Llama 3.3 70B Instructmath·Qwen2.5 72B Instructreasoning·Llama 3.1 70B Instructagentic·Llama 3.1 70B Instruct
Hype vs Reality
Llama 3.1 70B Instruct
#216 by perf·#18 by attention
QUIET
Llama 3.3 70B Instruct
#218 by perf·#18 by attention
QUIET
Qwen2.5 72B Instruct
#128 by perf·#2 by attention
DESERVED
Best value
1.3x better value than Qwen2.5 72B Instruct
Llama 3.1 70B Instruct
90.7 pts/$
$0.40/M
Llama 3.3 70B Instruct
171.9 pts/$
$0.21/M
Qwen2.5 72B Instruct
129.5 pts/$
$0.38/M
Vendor risk
Meta logo
Meta AI
$1.87T·Tier 1
Low risk
Meta logo
Meta AI
$1.87T·Tier 1
Low risk
Alibaba Qwen logo
Alibaba (Qwen)
$293.0B·Tier 1
Low risk
Head to head
Llama 3.1 70B InstructLlama 3.3 70B InstructQwen2.5 72B Instruct
Aider · Code Editing
Qwen2.5 72B Instruct leads by +6.0
Llama 3.1 70B Instruct
58.6
Llama 3.3 70B Instruct
59.4
Qwen2.5 72B Instruct
65.4
Chatbot Arena Elo · Overall
Llama 3.3 70B Instruct leads by +15.0
Llama 3.1 70B Instruct
1293.3
Llama 3.3 70B Instruct
1317.8
Qwen2.5 72B Instruct
1302.8
Balrog
Llama 3.1 70B Instruct leads by +4.9
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
Llama 3.1 70B Instruct
27.9
Llama 3.3 70B Instruct
23.0
Qwen2.5 72B Instruct
16.2
Dtbench
Qwen2.5 72B Instruct leads by +4.9
Llama 3.1 70B Instruct
33.3
Llama 3.3 70B Instruct
32.5
Qwen2.5 72B Instruct
38.2
GPQA diamond
Qwen2.5 72B Instruct leads by +2.3
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Llama 3.1 70B Instruct
25.6
Llama 3.3 70B Instruct
29.9
Qwen2.5 72B Instruct
32.2
BBH (HuggingFace)
Qwen2.5 72B Instruct leads by +5.3
Llama 3.1 70B Instruct
55.9
Llama 3.3 70B Instruct
56.6
Qwen2.5 72B Instruct
61.9
GPQA
Qwen2.5 72B Instruct leads by +2.5
Llama 3.1 70B Instruct
14.2
Llama 3.3 70B Instruct
10.5
Qwen2.5 72B Instruct
16.7
IFEval
Llama 3.3 70B Instruct leads by +3.3
Llama 3.1 70B Instruct
86.7
Llama 3.3 70B Instruct
90.0
Qwen2.5 72B Instruct
86.4
MATH Level 5
Qwen2.5 72B Instruct leads by +11.5
Llama 3.1 70B Instruct
38.1
Llama 3.3 70B Instruct
48.3
Qwen2.5 72B Instruct
59.8
MMLU-PRO
Qwen2.5 72B Instruct leads by +3.3
Llama 3.1 70B Instruct
47.9
Llama 3.3 70B Instruct
48.1
Qwen2.5 72B Instruct
51.4
MUSR
Llama 3.1 70B Instruct leads by +2.1
Llama 3.1 70B Instruct
17.7
Llama 3.3 70B Instruct
15.6
Qwen2.5 72B Instruct
11.7
Lmca
Llama 3.3 70B Instruct leads by +3.1
Llama 3.1 70B Instruct
17.5
Llama 3.3 70B Instruct
20.6
Qwen2.5 72B Instruct
15.8
MATH level 5
Qwen2.5 72B Instruct leads by +21.6
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Llama 3.1 70B Instruct
36.7
Llama 3.3 70B Instruct
41.6
Qwen2.5 72B Instruct
63.2
MMLU
Llama 3.3 70B Instruct leads by +1.3
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Llama 3.1 70B Instruct
73.5
Llama 3.3 70B Instruct
81.7
Qwen2.5 72B Instruct
80.4
OTIS Mock AIME 2024-2025
Qwen2.5 72B Instruct leads by +2.9
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Llama 3.1 70B Instruct
3.5
Llama 3.3 70B Instruct
5.0
Qwen2.5 72B Instruct
8.0
WeirdML
Qwen2.5 72B Instruct leads by +1.5
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Llama 3.1 70B Instruct
9.0
Llama 3.3 70B Instruct
14.4
Qwen2.5 72B Instruct
16.0
CMMLU
Qwen2.5 72B Instruct leads by +21.3
Llama 3.1 70B Instruct
64.4
Qwen2.5 72B Instruct
85.7
The Agent Company
Llama 3.1 70B Instruct leads by +1.2
The Agent Company · tests AI agents on realistic corporate tasks like email management, code review, data analysis, and cross-tool workflows.
Llama 3.1 70B Instruct
6.9
Qwen2.5 72B Instruct
5.7
Full benchmark table
BenchmarkLlama 3.1 70B InstructLlama 3.3 70B InstructQwen2.5 72B Instruct
Aider · Code Editing
58.659.465.4
Chatbot Arena Elo · Overall
1293.31317.81302.8
Balrog
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
27.923.016.2
Dtbench
33.332.538.2
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
25.629.932.2
BBH (HuggingFace)
55.956.661.9
GPQA
14.210.516.7
IFEval
86.790.086.4
MATH Level 5
38.148.359.8
MMLU-PRO
47.948.151.4
MUSR
17.715.611.7
Lmca
17.520.615.8
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
36.741.663.2
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
73.581.780.4
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
3.55.08.0
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
9.014.416.0
CMMLU
64.4—85.7
The Agent Company
The Agent Company · tests AI agents on realistic corporate tasks like email management, code review, data analysis, and cross-tool workflows.
6.9—5.7
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Meta logoLlama 3.1 70B Instruct$0.40$0.40131K tokens (~66 books)$4.00
Meta logoLlama 3.3 70B Instruct$0.10$0.32131K tokens (~66 books)$1.55
Alibaba Qwen logoQwen2.5 72B Instruct$0.36$0.4033K tokens (~16 books)$3.70