Tested on 4 benchmarks · BenchGecko score 17.0. Top scores: GPQA diamond (31.2%), OTIS Mock AIME 2024-2025 (29.9%), ARC-AGI (5.0%).
Abstraction and Reasoning Corpus. Tests fluid intelligence through novel visual pattern recognition puzzles. Core measure of general intelligence.
ARC-AGI 2, harder sequel to ARC. More complex abstract reasoning patterns that test generalization ability beyond training data.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
Graduate-level science questions written by PhD experts. Diamond subset contains questions where experts disagree, testing deep understanding.
- Typetext
- ContextN/A
- ReleasedJan 2024
- LicenseProprietary
- Statusbenchmark-only
Frequently Asked Questions
Key facts · as of 2026-04-09
- Magistral Small 1.1 by Unknown. BenchGecko score 17.0, rank 289 of 312 scored models (normalized average of public benchmark scores).
- List price n/a input · n/a output per 1M tokens (as of 2026-04-09).
How to cite · data as of 2026-04-09
Magistral Small 1.1 · benchmarks, pricing and providers. BenchGecko, data as of 2026-04-09. https://benchgecko.ai/model/magistral-small-1-1
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP