Compare · ModelsLive · 2 picked · head to head

Gemma 2 27B vs Llama 3.1 70B Instruct

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Llama 3.1 70B Instruct wins 12 of 13 shared benchmarks. Leads in arena · general · knowledge.

Category leads
arena·Llama 3.1 70B Instructgeneral·Llama 3.1 70B Instructknowledge·Llama 3.1 70B Instructlanguage·Llama 3.1 70B Instructmath·Llama 3.1 70B Instructreasoning·Llama 3.1 70B Instruct
Hype vs Reality
Gemma 2 27B
#220 by perf·#7 by attention
OVERHYPED
Llama 3.1 70B Instruct
#216 by perf·#18 by attention
QUIET
Best value
1.6x better value than Gemma 2 27B
Gemma 2 27B
55.2 pts/$
$0.65/M
Llama 3.1 70B Instruct
90.7 pts/$
$0.40/M
Vendor risk
Google DeepMind logo
Google DeepMind
$4.20T·Tier 1
Low risk
Meta logo
Meta AI
$1.87T·Tier 1
Low risk
Head to head
Gemma 2 27BLlama 3.1 70B Instruct
Chatbot Arena Elo · Overall
Llama 3.1 70B Instruct leads by +3.9
Gemma 2 27B
1289.4
Llama 3.1 70B Instruct
1293.3
Dtbench
Llama 3.1 70B Instruct leads by +20.0
Gemma 2 27B
13.3
Llama 3.1 70B Instruct
33.3
GPQA diamond
Llama 3.1 70B Instruct leads by +10.3
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Gemma 2 27B
15.3
Llama 3.1 70B Instruct
25.6
BBH (HuggingFace)
Llama 3.1 70B Instruct leads by +6.7
Gemma 2 27B
49.3
Llama 3.1 70B Instruct
55.9
GPQA
Gemma 2 27B leads by +2.5
Gemma 2 27B
16.7
Llama 3.1 70B Instruct
14.2
IFEval
Llama 3.1 70B Instruct leads by +6.9
Gemma 2 27B
79.8
Llama 3.1 70B Instruct
86.7
MATH Level 5
Llama 3.1 70B Instruct leads by +14.2
Gemma 2 27B
23.9
Llama 3.1 70B Instruct
38.1
MMLU-PRO
Llama 3.1 70B Instruct leads by +9.5
Gemma 2 27B
38.4
Llama 3.1 70B Instruct
47.9
MUSR
Llama 3.1 70B Instruct leads by +8.6
Gemma 2 27B
9.1
Llama 3.1 70B Instruct
17.7
Lmca
Llama 3.1 70B Instruct leads by +9.1
Gemma 2 27B
8.3
Llama 3.1 70B Instruct
17.5
MATH level 5
Llama 3.1 70B Instruct leads by +8.8
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Gemma 2 27B
27.9
Llama 3.1 70B Instruct
36.7
MMLU
Llama 3.1 70B Instruct leads by +5.9
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Gemma 2 27B
67.6
Llama 3.1 70B Instruct
73.5
OTIS Mock AIME 2024-2025
Llama 3.1 70B Instruct leads by +2.2
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Gemma 2 27B
1.3
Llama 3.1 70B Instruct
3.5
Full benchmark table
BenchmarkGemma 2 27BLlama 3.1 70B Instruct
Chatbot Arena Elo · Overall
1289.41293.3
Dtbench
13.333.3
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
15.325.6
BBH (HuggingFace)
49.355.9
GPQA
16.714.2
IFEval
79.886.7
MATH Level 5
23.938.1
MMLU-PRO
38.447.9
MUSR
9.117.7
Lmca
8.317.5
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
27.936.7
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
67.673.5
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
1.33.5
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Google DeepMind logoGemma 2 27B$0.65$0.658K tokens (~4 books)$6.50
Meta logoLlama 3.1 70B Instruct$0.40$0.40131K tokens (~66 books)$4.00