Live11 categories · 1,235 models tracked

Models by Category.
Ranked.

Every model category we track · one leaderboard per task. Pick the skill you care about and see which models dominate it.

104 models

AI models ranked by coding benchmarks. Compare HumanEval+, SWE-bench Verified, Aider Polyglot, and more across all providers.

Top 3
1GPT-5.5
85.0
2GPT-5 Chat
81.9
3Claude Mythos Preview
81.8
225 models

AI models ranked by reasoning benchmarks. Compare GPQA Diamond, ARC-AGI, BBH, and other reasoning tests across all providers.

Top 3
1GPT-5.5 Pro
87.8
2GPT-5.5
85.0
3Claude Opus 5.5
82.2
202 models

AI models ranked by math benchmarks. Compare MATH-500, GSM8K, and competition-level math scores across all providers.

Top 3
1Claude Opus 5.5
82.2
2GPT-6 Astra
76.7
3Claude Sonnet 5.5
75.0
216 models

AI models ranked by knowledge benchmarks. Compare MMLU-Pro, GPQA Diamond, SimpleQA, and other knowledge tests.

Top 3
1GPT-5.5 Pro
87.8
2GPT-5.5
85.0
3Claude Opus 5.5
82.2
27 models

AI models ranked by vision and multimodal benchmarks. Compare MMMU, VideoMME, and visual reasoning scores.

Top 3
1Gemini 3 Pro
56.1
2GPT-5
52.9
3GPT-5 Mini
51.5
177 models

Open-source AI models ranked by benchmark score. Compare Llama, Mistral, DeepSeek, Qwen, and other open-weight models.

Top 3
1DeepSeek V3.2 Speciale
78.2
2Step 3.5 Flash
76.9
3DeepSeek-V2 (MoE-236B, May 2024)
76.5
124 models

Cheapest AI models ranked by score. All models with input pricing under $1 per million tokens, sorted by benchmark performance.

Top 3
1DeepSeek V3.2 Speciale
78.2
2Step 3.5 Flash
76.9
3MiMo-V2-Flash
73.3
176 models

AI models under $5 per million tokens ranked by benchmark score. The sweet spot of price and performance.

Top 3
1Claude Opus 5.5
82.2
2GPT-5 Chat
81.9
3DeepSeek V3.2 Speciale
78.2
62 models

AI models with 1 million+ token context windows ranked by score. Compare Gemini, Claude, and other long-context models.

Top 3
1Claude Opus 5.5
82.2
2Claude Mythos Preview
81.8
3Gemini 2.5 Pro Preview 05-06
76.9
106 models

Small AI models (under 10B parameters) ranked by benchmark score. Lightweight models you can run locally.

Top 3
1Gemini 2.5 Pro Preview 05-06
76.9
2o4 Mini High
72.0
3MiniMax M2
69.5
320 models

The best AI model from each provider, ranked by benchmark score. Compare the flagships from OpenAI, Anthropic, Google, Meta, and more.

Top 3
1GPT-5.5 Pro
87.8
2GPT-5.5
85.0
3Claude Opus 5.5
82.2