Benchmarks
How models are measured · SWE-bench, GPQA, MMLU.
Top 12 terms
The real-world coding benchmark · AI resolves actual GitHub issues in open-source Python repos.
A 164-problem Python benchmark where the model writes a function from its docstring and passes unit tests.
A contamination-resistant benchmark that refreshes tasks monthly to prevent models from memorizing answers.
Crowdsourced head-to-head AI model comparison · humans vote on anonymous outputs and Elo ratings rank the models.
A harder version of MMLU with 10 answer choices, filtered noise, and more reasoning-heavy questions.
ARC AI2 · knowledge benchmark tracked on BenchGecko.
BBH · reasoning benchmark tracked on BenchGecko.
GSM8K · math benchmark tracked on BenchGecko.
HellaSwag · knowledge benchmark tracked on BenchGecko.
LAMBADA · knowledge benchmark tracked on BenchGecko.
MMLU · knowledge benchmark tracked on BenchGecko.
GPQA diamond · knowledge benchmark tracked on BenchGecko.