Chatbot Arena Elo · Coding
The Frontier
Best score over time · one chart, every benchmark
Full rankings
58 models tested · sorted by score
Score distribution
Where models cluster
Correlated benchmarks
Pearson r · original research
Benchmarks that track with Chatbot Arena Elo · Coding
Pearson correlation across models scored on both benchmarks. Closer to 1 = strongly predictive.
Frequently asked
About Chatbot Arena Elo · Coding
What does Chatbot Arena Elo · Coding measure?
Chatbot Arena Elo · Coding is a knowledge benchmark in the BenchGecko catalog. 58 AI models have been tested on it. Scores range from 1170.0 to 1671.2 out of 1600.
Which model leads on Chatbot Arena Elo · Coding?
Qwen3.8 Max (0902) from Alibaba Qwen leads Chatbot Arena Elo · Coding with a score of 1671.2. The median score across 58 tested models is 1433.7.
Is Chatbot Arena Elo · Coding saturated?
Yes · the top model on Chatbot Arena Elo · Coding has reached 1671.2 out of 1600, within 5% of the theoretical ceiling. This benchmark is approaching saturation and may be replaced by a harder successor.
Does Chatbot Arena Elo · Coding predict performance on other benchmarks?
Yes · Chatbot Arena Elo · Coding scores correlate 0.93 with Metr Time Horizons across 7 shared models. Models that do well on Chatbot Arena Elo · Coding tend to do well on Metr Time Horizons.
How often is Chatbot Arena Elo · Coding data refreshed?
BenchGecko pulls updates daily. New model scores on Chatbot Arena Elo · Coding appear as soon as they are published by Epoch AI or the model provider.
- Category
- Knowledge
- Max score
- 1600
- Models
- 58
- Updated
- 2026-09-21
Top on Chatbot Arena Elo · Coding
Qwen3.8 Max (0902) · 1671.2Claude Fable 5 · 1653.9MiMo-V2.6-Pro · 1618.4GLM 5.3 Flash · 1615.6GLM 5.2 · 1593.3More knowledge benchmarks
Same category · related evaluations