Compare · ModelsLive · 2 picked · head to head
Llama 3.1 70B Instruct vs Llama 3.3 70B Instruct
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Llama 3.3 70B Instruct wins on 12/16 benchmarks
Llama 3.3 70B Instruct wins 12 of 16 shared benchmarks. Leads in coding · arena · knowledge.
Category leads
coding·Llama 3.3 70B Instructarena·Llama 3.3 70B Instructknowledge·Llama 3.3 70B Instructgeneral·Llama 3.3 70B Instructlanguage·Llama 3.3 70B Instructmath·Llama 3.3 70B Instructreasoning·Llama 3.1 70B Instruct
Hype vs Reality
Attention vs performance
Llama 3.1 70B Instruct
#216 by perf·#18 by attention
Llama 3.3 70B Instruct
#218 by perf·#18 by attention
Best value
Llama 3.3 70B Instruct
1.9x better value than Llama 3.1 70B Instruct
Llama 3.1 70B Instruct
90.7 pts/$
$0.40/M
Llama 3.3 70B Instruct
171.9 pts/$
$0.21/M
Vendor risk
Who is behind the model
Meta AI
$1.87T·Tier 1
Meta AI
$1.87T·Tier 1
Head to head
16 benchmarks · 2 models
Llama 3.1 70B InstructLlama 3.3 70B Instruct
Aider · Code Editing
Llama 3.3 70B Instruct leads by +0.8
Llama 3.1 70B Instruct
58.6
Llama 3.3 70B Instruct
59.4
Chatbot Arena Elo · Overall
Llama 3.3 70B Instruct leads by +24.5
Llama 3.1 70B Instruct
1293.3
Llama 3.3 70B Instruct
1317.8
Balrog
Llama 3.1 70B Instruct leads by +4.9
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
Llama 3.1 70B Instruct
27.9
Llama 3.3 70B Instruct
23.0
Dtbench
Llama 3.1 70B Instruct leads by +0.9
Llama 3.1 70B Instruct
33.3
Llama 3.3 70B Instruct
32.5
GPQA diamond
Llama 3.3 70B Instruct leads by +4.3
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Llama 3.1 70B Instruct
25.6
Llama 3.3 70B Instruct
29.9
BBH (HuggingFace)
Llama 3.3 70B Instruct leads by +0.6
Llama 3.1 70B Instruct
55.9
Llama 3.3 70B Instruct
56.6
GPQA
Llama 3.1 70B Instruct leads by +3.7
Llama 3.1 70B Instruct
14.2
Llama 3.3 70B Instruct
10.5
IFEval
Llama 3.3 70B Instruct leads by +3.3
Llama 3.1 70B Instruct
86.7
Llama 3.3 70B Instruct
90.0
MATH Level 5
Llama 3.3 70B Instruct leads by +10.3
Llama 3.1 70B Instruct
38.1
Llama 3.3 70B Instruct
48.3
MMLU-PRO
Llama 3.3 70B Instruct leads by +0.3
Llama 3.1 70B Instruct
47.9
Llama 3.3 70B Instruct
48.1
MUSR
Llama 3.1 70B Instruct leads by +2.1
Llama 3.1 70B Instruct
17.7
Llama 3.3 70B Instruct
15.6
Lmca
Llama 3.3 70B Instruct leads by +3.1
Llama 3.1 70B Instruct
17.5
Llama 3.3 70B Instruct
20.6
MATH level 5
Llama 3.3 70B Instruct leads by +4.9
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Llama 3.1 70B Instruct
36.7
Llama 3.3 70B Instruct
41.6
MMLU
Llama 3.3 70B Instruct leads by +8.3
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Llama 3.1 70B Instruct
73.5
Llama 3.3 70B Instruct
81.7
OTIS Mock AIME 2024-2025
Llama 3.3 70B Instruct leads by +1.5
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Llama 3.1 70B Instruct
3.5
Llama 3.3 70B Instruct
5.0
WeirdML
Llama 3.3 70B Instruct leads by +5.5
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Llama 3.1 70B Instruct
9.0
Llama 3.3 70B Instruct
14.4
Full benchmark table
| Benchmark | Llama 3.1 70B Instruct | Llama 3.3 70B Instruct |
|---|---|---|
Aider · Code Editing | 58.6 | 59.4 |
Chatbot Arena Elo · Overall | 1293.3 | 1317.8 |
Balrog Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning. | 27.9 | 23.0 |
Dtbench | 33.3 | 32.5 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 25.6 | 29.9 |
BBH (HuggingFace) | 55.9 | 56.6 |
GPQA | 14.2 | 10.5 |
IFEval | 86.7 | 90.0 |
MATH Level 5 | 38.1 | 48.3 |
MMLU-PRO | 47.9 | 48.1 |
MUSR | 17.7 | 15.6 |
Lmca | 17.5 | 20.6 |
MATH level 5 MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics. | 36.7 | 41.6 |
MMLU Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge. | 73.5 | 81.7 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 3.5 | 5.0 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 9.0 | 14.4 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.40 | $0.40 | 131K tokens (~66 books) | $4.00 | |
| $0.10 | $0.32 | 131K tokens (~66 books) | $1.55 |
People also compared
GPT-5.5 Pro vs Llama 3.1 70B InstructGPT-5.5 vs Llama 3.1 70B InstructClaude Opus 5.5 vs Llama 3.1 70B InstructClaude Mythos Preview vs Llama 3.1 70B InstructDeepSeek V3.2 Speciale vs Llama 3.1 70B InstructDeepSeek-V2 (MoE-236B, May 2024) vs Llama 3.1 70B InstructGrok 3 Beta vs Llama 3.1 70B InstructLlama 3.1 70B Instruct vs MiniMax M2