MMLU
MMLU is a knowledge benchmark tracked on BenchGecko. Massive Multitask Language Understanding. 57 subjects from STEM, humanities, and social sciences. The most widely-cited knowledge benchmark.
Built from BenchGecko data as of October 2, 2026 · updates when the data changes
MMLU is a knowledge benchmark tracked on BenchGecko. Massive Multitask Language Understanding. 57 subjects from STEM, humanities, and social sciences. The most widely-cited knowledge benchmark.
Basic
Massive Multitask Language Understanding. 57 subjects from STEM, humanities, and social sciences. The most widely-cited knowledge benchmark.
Deep
Scores are reported (%, maximum 100); higher is better. BenchGecko collects MMLU scores from Epoch AI and refreshes them daily; the leaderboard on this page shows the current top models.
Expert
Compare models only on the same benchmark version and settings. A benchmark separates models well while scores are spread out; when top models bunch together near the maximum, it stops telling them apart.
Depending on why you're here
- ·knowledge benchmark (%, maximum 100)
- ·Source: Epoch AI
- ·Useful if your workload is knowledge
- ·Check the live leaderboard on /benchmark/mmlu before choosing a model
- ·Labs cite benchmark results in launch announcements; check the source and version
- ·MMLU is a test of how good an AI is at knowledge
- ·A higher score means better results on that kind of task