Compare · ModelsLive · 2 picked · head to head

Gemma 4 26B A4B vs Llama 3.3 70B Instruct

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Gemma 4 26B A4B wins 6 of 6 shared benchmarks. Leads in arena · general · knowledge.

Category leads
arena·Gemma 4 26B A4B general·Gemma 4 26B A4B knowledge·Gemma 4 26B A4B math·Gemma 4 26B A4B coding·Gemma 4 26B A4B
Hype vs Reality
Gemma 4 26B A4B
#151 by perf·#7 by attention
DESERVED
Llama 3.3 70B Instruct
#218 by perf·#18 by attention
QUIET
Best value
1.8x better value than Llama 3.3 70B Instruct
Gemma 4 26B A4B
314.5 pts/$
$0.15/M
Llama 3.3 70B Instruct
171.9 pts/$
$0.21/M
Vendor risk
Google DeepMind logo
Google DeepMind
$4.20T·Tier 1
Low risk
Meta logo
Meta AI
$1.87T·Tier 1
Low risk
Head to head
Gemma 4 26B A4B Llama 3.3 70B Instruct
Chatbot Arena Elo · Overall
Gemma 4 26B A4B leads by +119.7
Gemma 4 26B A4B
1437.5
Llama 3.3 70B Instruct
1317.8
Dtbench
Gemma 4 26B A4B leads by +25.8
Gemma 4 26B A4B
58.2
Llama 3.3 70B Instruct
32.5
GPQA diamond
Gemma 4 26B A4B leads by +34.4
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Gemma 4 26B A4B
64.3
Llama 3.3 70B Instruct
29.9
Lmca
Gemma 4 26B A4B leads by +14.3
Gemma 4 26B A4B
34.9
Llama 3.3 70B Instruct
20.6
OTIS Mock AIME 2024-2025
Gemma 4 26B A4B leads by +77.2
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Gemma 4 26B A4B
82.2
Llama 3.3 70B Instruct
5.0
WeirdML
Gemma 4 26B A4B leads by +20.7
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Gemma 4 26B A4B
35.2
Llama 3.3 70B Instruct
14.4
Full benchmark table
BenchmarkGemma 4 26B A4B Llama 3.3 70B Instruct
Chatbot Arena Elo · Overall
1437.51317.8
Dtbench
58.232.5
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
64.329.9
Lmca
34.920.6
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
82.25.0
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
35.214.4
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Google DeepMind logoGemma 4 26B A4B $0.07$0.23262K tokens (~131 books)$1.07
Meta logoLlama 3.3 70B Instruct$0.10$0.32131K tokens (~66 books)$1.55