Gemma 2 9B by Google is an advanced, open-source language model that sets a new standard for efficiency and performance in its size class. Designed for a wide variety of...
Tested on 13 benchmarks · BenchGecko score 51.1. Top scores: Chatbot Arena Elo — Overall (1265.0%), GSM8K (84.9%), IFEval (74.4%).
HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.
Grade school math word problems. 8,500 problems testing multi-step arithmetic reasoning. A foundational math benchmark.
Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.
HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.
Physical Intuition QA. Tests understanding of everyday physical interactions and commonsense physics.
Massive Multitask Language Understanding. 57 subjects from STEM, humanities, and social sciences. The most widely-cited knowledge benchmark.
HuggingFace MMLU-Pro. Harder version of MMLU with 10 answer choices instead of 4 and more challenging questions.
- Typetext
- Context8K tokens (~4 books)
- ReleasedJun 2024
- LicenseOpen Source
- StatusActive
- Cost / Message~$0.000
Frequently Asked Questions
Key facts · as of 2026-04-14
- Gemma 2 9B by Google DeepMind. BenchGecko score 51.1, rank 138 of 312 scored models (normalized average of public benchmark scores).
- List price $0.0300 input · $0.0900 output per 1M tokens (as of 2026-04-14).
How to cite · data as of 2026-04-14
Gemma 2 9B · benchmarks, pricing and providers. BenchGecko, data as of 2026-04-14. https://benchgecko.ai/model/gemma-2-9b-it
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP