ARC-AGI-2
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
67 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with ARC-AGI-2
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About ARC-AGI-2
What does ARC-AGI-2 measure?
ARC-AGI-2 · the second iteration of the Abstraction and Reasoning Corpus, testing novel pattern recognition and abstract reasoning without prior training data. 67 AI models have been tested on it. Scores range from 0.1 to 95.0 out of 100.
Which model leads on ARC-AGI-2?
GPT-6 Astra from OpenAI leads ARC-AGI-2 with a score of 95.0. The median score across 67 tested models is 16.0.
Is ARC-AGI-2 saturated?
No · the top score is 95.0 out of 100 (95%). There is still meaningful room for improvement on ARC-AGI-2.
Does ARC-AGI-2 predict performance on other benchmarks?
Yes · ARC-AGI-2 scores correlate 0.96 with Artificial Analysis · GPQA Diamond across 7 shared models. Models that do well on ARC-AGI-2 tend to do well on Artificial Analysis · GPQA Diamond.
How often is ARC-AGI-2 data refreshed?
BenchGecko pulls updates daily. New model scores on ARC-AGI-2 appear as soon as they are published by Epoch AI or the model provider.
- Category
- Reasoning
- Max score
- 100
- Models
- 67
- Updated
- 2026-09-22
Top on ARC-AGI-2
GPT-6 Astra · 95.0Claude Opus 5.5 · 92.5GPT-5.6 Sol · 92.5Claude Opus 5 · 90.4Claude Fable 5.1 · 90.0More reasoning benchmarks
Same category · related evaluations