Balrog
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
36 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 68.3 | |
| 2 | 63.4 | |
| 3 | 60.0 | |
| 4 | 58.1 | |
| 5 | 57.0 | |
| 6 | 53.2 | |
| 7 | 48.1 | |
| 8 | 45.6 | |
| 9 | 43.6 | |
| 10 | 43.5 | |
| 11 | 43.3 | |
| 12 | 34.9 | |
| 13 | 33.5 | |
| 14 | 32.8 | |
| 15 | 32.6 | |
| 16 | 32.3 | |
| 17 | 32.3 | |
| 18 | 31.2 | |
| 19 | 29.5 | |
| 20 | 27.9 | |
| 21 | 27.3 | |
| 22 | 23.0 | |
| 23 | 23.0 | |
| 24 | 21.0 | |
| 25 | 21.0 | |
| 26 | 19.5 | |
| 27 | 19.3 | |
| 28 | 17.6 | |
| 29 | 17.4 | |
| 30 | 17.4 | |
| 31 | 16.2 | |
| 32 | 15.1 | |
| 33 | 14.6 | |
| 34 | 11.6 | |
| 35 | 7.8 | |
| 36 | 6.6 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Balrog
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Balrog
What does Balrog measure?
Balrog · benchmarks AI agents on text-based adventure games, testing language understanding, strategic planning, and long-horizon reasoning. 36 AI models have been tested on it. Scores range from 6.6 to 68.3 out of 100.
Which model leads on Balrog?
GPT-6 Astra from OpenAI leads Balrog with a score of 68.3. The median score across 36 tested models is 30.4.
Is Balrog saturated?
No · the top score is 68.3 out of 100 (68%). There is still meaningful room for improvement on Balrog.
Does Balrog predict performance on other benchmarks?
Yes · Balrog scores correlate 0.96 with Lmca across 19 shared models. Models that do well on Balrog tend to do well on Lmca.
How often is Balrog data refreshed?
BenchGecko pulls updates daily. New model scores on Balrog appear as soon as they are published by Epoch AI or the model provider.
- Category
- Reasoning
- Max score
- 100
- Models
- 36
- Updated
- 2026-09-04
Top on Balrog
GPT-6 Astra · 68.3Claude Opus 5 · 63.4GPT-5.6 Sol · 60.0Gemini 3 Pro · 58.1Gemini 3.1 Pro Preview · 57.0More reasoning benchmarks
Same category · related evaluations