Tested on 16 benchmarks · BenchGecko score 47.7. Top scores: HELM — IFEval (95.1%), MATH level 5 (90.9%), HELM — MMLU-Pro (79.9%).
Multi-language code editing from Aider. Tests editing ability across Python, JavaScript, TypeScript, Java, C++, Go, Rust, and more.
Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.
Stanford HELM WildBench evaluation. Tests reasoning on challenging real-world tasks.
Abstraction and Reasoning Corpus. Tests fluid intelligence through novel visual pattern recognition puzzles. Core measure of general intelligence.
ARC-AGI 2, harder sequel to ARC. More complex abstract reasoning patterns that test generalization ability beyond training data.
Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
Stanford HELM evaluation of mathematical reasoning across diverse problem types.
- Typetext
- ContextN/A
- ReleasedJan 2024
- LicenseProprietary
- Statusbenchmark-only
Frequently Asked Questions
Key facts · as of 2026-05-03
- Grok-3 mini by xAI. BenchGecko score 47.7, rank 160 of 312 scored models (normalized average of public benchmark scores).
- List price n/a input · n/a output per 1M tokens (as of 2026-05-03).
How to cite · data as of 2026-05-03
Grok-3 mini · benchmarks, pricing and providers. BenchGecko, data as of 2026-05-03. https://benchgecko.ai/model/grok-3-mini
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP