Compare · ModelsLive · 2 picked · head to head

Gemma 2 27B vs Mistral 7B V0.1

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Gemma 2 27B wins 8 of 9 shared benchmarks. Leads in math · general · knowledge.

Category leads
math·Gemma 2 27Bgeneral·Gemma 2 27Bknowledge·Gemma 2 27Blanguage·Gemma 2 27Breasoning·Mistral 7B V0.1
Hype vs Reality
Gemma 2 27B
#220 by perf·#7 by attention
OVERHYPED
Mistral 7B V0.1
#176 by perf·no signal
QUIET
Best value
Gemma 2 27B
55.2 pts/$
$0.65/M
Mistral 7B V0.1
n/a
no price
Vendor risk
Google DeepMind logo
Google DeepMind
$4.20T·Tier 1
Low risk
Mistral AI logo
Mistral AI
$14.0B·Tier 1
Medium risk
Head to head
Gemma 2 27BMistral 7B V0.1
GSM8K
Gemma 2 27B leads by +30.5
Grade School Math 8K · 8,500 linguistically diverse grade-school math word problems that require multi-step reasoning to solve.
Gemma 2 27B
84.9
Mistral 7B V0.1
54.4
BBH (HuggingFace)
Gemma 2 27B leads by +27.3
Gemma 2 27B
49.3
Mistral 7B V0.1
22.0
GPQA
Gemma 2 27B leads by +11.1
Gemma 2 27B
16.7
Mistral 7B V0.1
5.6
IFEval
Gemma 2 27B leads by +55.9
Gemma 2 27B
79.8
Mistral 7B V0.1
23.9
MATH Level 5
Gemma 2 27B leads by +20.9
Gemma 2 27B
23.9
Mistral 7B V0.1
3.0
MMLU-PRO
Gemma 2 27B leads by +16.0
Gemma 2 27B
38.4
Mistral 7B V0.1
22.4
MUSR
Mistral 7B V0.1 leads by +1.6
Gemma 2 27B
9.1
Mistral 7B V0.1
10.7
MMLU
Gemma 2 27B leads by +17.6
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Gemma 2 27B
67.6
Mistral 7B V0.1
50.0
PIQA
Gemma 2 27B leads by +1.4
PIQA (Physical Interaction QA) · tests intuitive physical reasoning by asking models to select the correct approach for everyday physical tasks.
Gemma 2 27B
67.4
Mistral 7B V0.1
66.0
Full benchmark table
BenchmarkGemma 2 27BMistral 7B V0.1
GSM8K
Grade School Math 8K · 8,500 linguistically diverse grade-school math word problems that require multi-step reasoning to solve.
84.954.4
BBH (HuggingFace)
49.322.0
GPQA
16.75.6
IFEval
79.823.9
MATH Level 5
23.93.0
MMLU-PRO
38.422.4
MUSR
9.110.7
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
67.650.0
PIQA
PIQA (Physical Interaction QA) · tests intuitive physical reasoning by asking models to select the correct approach for everyday physical tasks.
67.466.0
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Google DeepMind logoGemma 2 27B$0.65$0.658K tokens (~4 books)$6.50
Mistral AI logoMistral 7B V0.1————