Compare · ModelsLive · 2 picked · head to head

Mistral Large vs o3 Mini

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

o3 Mini wins 6 of 7 shared benchmarks. Leads in general · math · knowledge.

Category leads
general·o3 Minimath·o3 Miniknowledge·o3 Minireasoning·o3 Mini
Hype vs Reality
Mistral Large
#185 by perf·#19 by attention
QUIET
o3 Mini
#239 by perf·no signal
QUIET
Best value
1.2x better value than Mistral Large
Mistral Large
10.2 pts/$
$4.00/M
o3 Mini
11.8 pts/$
$2.75/M
Vendor risk
Mistral AI logo
Mistral AI
$14.0B·Tier 1
Medium risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Head to head
Mistral Largeo3 Mini
Dtbench
o3 Mini leads by +21.5
Mistral Large
26.5
o3 Mini
48.0
FrontierMath-2025-02-28-Private
o3 Mini leads by +11.8
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Mistral Large
0.6
o3 Mini
12.4
GPQA diamond
o3 Mini leads by +51.0
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Mistral Large
18.4
o3 Mini
69.4
Lech Mazur Writing
Mistral Large leads by +7.3
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
Mistral Large
69.0
o3 Mini
61.7
MATH level 5
o3 Mini leads by +72.0
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Mistral Large
24.5
o3 Mini
96.5
OTIS Mock AIME 2024-2025
o3 Mini leads by +75.1
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Mistral Large
1.9
o3 Mini
76.9
SimpleBench
o3 Mini leads by +0.4
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Mistral Large
7.0
o3 Mini
7.4
Full benchmark table
BenchmarkMistral Largeo3 Mini
Dtbench
26.548.0
FrontierMath-2025-02-28-Private
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
0.612.4
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
18.469.4
Lech Mazur Writing
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
69.061.7
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
24.596.5
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
1.976.9
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
7.07.4
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Mistral AI logoMistral Large$2.00$6.00128K tokens (~64 books)$30.00
OpenAI logoo3 Mini$1.10$4.40200K tokens (~100 books)$19.25