DeepResearch Bench
DeepResearch Bench · evaluates AI on complex multi-step research tasks requiring information gathering, synthesis, and producing comprehensive analyses.
The Frontier
Best score over time · one chart, every benchmark
Full rankings
26 models tested · sorted by score
| # | Model | Score |
|---|---|---|
| 1 | 55.3 | |
| 2 | 54.9 | |
| 3 | 54.8 | |
| 4 | 52.6 | |
| 5 | 50.2 | |
| 6 | 49.8 | |
| 7 | 49.6 | |
| 8 | 48.3 | |
| 9 | 47.8 | |
| 10 | 47.8 | |
| 11 | 47.3 | |
| 12 | 46.8 | |
| 13 | 46.3 | |
| 14 | 45.5 | |
| 15 | 45.2 | |
| 16 | 43.6 | |
| 17 | 42.8 | |
| 18 | 42.8 | |
| 19 | 41.1 | |
| 20 | 37.3 | |
| 21 | 36.3 | |
| 22 | 35.1 | |
| 23 | 35.1 | |
| 24 | 35.1 | |
| 25 | 29.3 | |
| 26 | 29.2 |
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with DeepResearch Bench
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About DeepResearch Bench
What does DeepResearch Bench measure?
DeepResearch Bench · evaluates AI on complex multi-step research tasks requiring information gathering, synthesis, and producing comprehensive analyses. 26 AI models have been tested on it. Scores range from 29.2 to 55.3 out of 100.
Which model leads on DeepResearch Bench?
Claude Opus 4.6 from Anthropic leads DeepResearch Bench with a score of 55.3. The median score across 26 tested models is 45.9.
Is DeepResearch Bench saturated?
No · the top score is 55.3 out of 100 (55%). There is still meaningful room for improvement on DeepResearch Bench.
Does DeepResearch Bench predict performance on other benchmarks?
Yes · DeepResearch Bench scores correlate 0.97 with Cybench across 8 shared models. Models that do well on DeepResearch Bench tend to do well on Cybench.
How often is DeepResearch Bench data refreshed?
BenchGecko pulls updates daily. New model scores on DeepResearch Bench appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 100
- Models
- 26
- Updated
- 2026-05-27
Top on DeepResearch Bench
Claude Opus 4.6 · 55.3Claude Sonnet 4.6 · 54.9Claude Opus 4.5 · 54.8Claude Sonnet 4.5 · 52.6Claude Opus 4.8 · 50.2More knowledge benchmarks
Same category · related evaluations