APEX-Agents
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
56 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with APEX-Agents
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About APEX-Agents
What does APEX-Agents measure?
APEX-Agents · evaluates AI agents on complex, multi-step tasks requiring planning, tool use, and autonomous decision-making in realistic environments. 56 AI models have been tested on it. Scores range from 1.1 to 75.5 out of 100.
Which model leads on APEX-Agents?
Claude Sonnet 5.5 from Anthropic leads APEX-Agents with a score of 75.5. The median score across 56 tested models is 33.0.
Is APEX-Agents saturated?
No · the top score is 75.5 out of 100 (76%). There is still meaningful room for improvement on APEX-Agents.
Does APEX-Agents predict performance on other benchmarks?
Yes · APEX-Agents scores correlate 0.93 with ARC-AGI-2 across 38 shared models. Models that do well on APEX-Agents tend to do well on ARC-AGI-2.
How often is APEX-Agents data refreshed?
BenchGecko pulls updates daily. New model scores on APEX-Agents appear as soon as they are published by Epoch AI or the model provider.
- Category
- Agent
- Max score
- 100
- Models
- 56
- Updated
- 2026-09-28
Top on APEX-Agents
Claude Sonnet 5.5 · 75.5Claude Opus 5.5 · 73.5Claude Fable 5.1 · 68.6Gemini 3.7 Flash · 67.8Claude Opus 5 · 65.8More agent benchmarks
Same category · related evaluations