Compare · ModelsLive · 2 picked · head to head
Claude Haiku 4.5 vs Mistral Large 2411
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Claude Haiku 4.5 wins on 4/4 benchmarks
Claude Haiku 4.5 wins 4 of 4 shared benchmarks. Leads in math · knowledge.
Category leads
math·Claude Haiku 4.5knowledge·Claude Haiku 4.5
Hype vs Reality
Attention vs performance
Claude Haiku 4.5
#231 by perf·#13 by attention
Mistral Large 2411
#153 by perf·#19 by attention
Best value
Mistral Large 2411
1.0x better value than Claude Haiku 4.5
Claude Haiku 4.5
11.3 pts/$
$3.00/M
Mistral Large 2411
11.4 pts/$
$4.00/M
Vendor risk
Who is behind the model
Anthropic
$965.0B·Tier 1
Mistral AI
$14.0B·Tier 1
Head to head
4 benchmarks · 2 models
Claude Haiku 4.5Mistral Large 2411
FrontierMath-2025-02-28-Private
Claude Haiku 4.5 leads by +10.0
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Claude Haiku 4.5
10.4
Mistral Large 2411
0.3
GPQA diamond
Claude Haiku 4.5 leads by +26.5
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Claude Haiku 4.5
61.6
Mistral Large 2411
35.1
MATH level 5
Claude Haiku 4.5 leads by +46.1
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Claude Haiku 4.5
96.4
Mistral Large 2411
50.3
OTIS Mock AIME 2024-2025
Claude Haiku 4.5 leads by +58.9
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Claude Haiku 4.5
66.6
Mistral Large 2411
7.7
Full benchmark table
| Benchmark | Claude Haiku 4.5 | Mistral Large 2411 |
|---|---|---|
FrontierMath-2025-02-28-Private FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning. | 10.4 | 0.3 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 61.6 | 35.1 |
MATH level 5 MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics. | 96.4 | 50.3 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 66.6 | 7.7 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $1.00 | $5.00 | 200K tokens (~100 books) | $20.00 | |
| $2.00 | $6.00 | 131K tokens (~66 books) | $30.00 |