Home/Models/Grok-3 mini
xAI logo

Grok-3 mini

by xAI · Released Jan 2024

47.7
avg score
Rank #160
Compare
Better than 49% of all models
Context
N/A
Input $/1M
n/a
Output $/1M
n/a
Type
text
License
Proprietary
Benchmarks
16 tested
Data as of
About

Tested on 16 benchmarks · BenchGecko score 47.7. Top scores: HELM — IFEval (95.1%), MATH level 5 (90.9%), HELM — MMLU-Pro (79.9%).

Capabilities
coding
45.9
#111 globally
reasoning
27.3
#129 globally
math
52.7
#112 globally
knowledge
62.8
#41 globally
language
95.1
#2 globally
Benchmark Scores
Compare All
Tested on 16 benchmarks · Ranked across 5 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
Aider polyglot

Multi-language code editing from Aider. Tests editing ability across Python, JavaScript, TypeScript, Java, C++, Go, Rust, and more.

49.3·
WeirdML

Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.

42.6·
HELM — WildBench

Stanford HELM WildBench evaluation. Tests reasoning on challenging real-world tasks.

65.1·
ARC-AGI

Abstraction and Reasoning Corpus. Tests fluid intelligence through novel visual pattern recognition puzzles. Core measure of general intelligence.

16.5·
ARC-AGI-2

ARC-AGI 2, harder sequel to ARC. More complex abstract reasoning patterns that test generalization ability beyond training data.

0.4·
MATH level 5

Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.

90.9·
OTIS Mock AIME 2024-2025

Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.

77.8·
HELM — Omni-MATH

Stanford HELM evaluation of mathematical reasoning across diverse problem types.

31.8·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Links
Documentation
Community
BenchGecko API
grok-3-mini
Specifications
  • Typetext
  • ContextN/A
  • ReleasedJan 2024
  • LicenseProprietary
  • Statusbenchmark-only
Available On
xAI logoxAIn/a
Share & Export
Tweet
Grok-3 mini is a proprietary text AI model by xAI, released in January 2024. It has an average benchmark score of 47.7.

Key facts · as of 2026-05-03

  • Grok-3 mini by xAI. BenchGecko score 47.7, rank 160 of 312 scored models (normalized average of public benchmark scores).
  • List price n/a input · n/a output per 1M tokens (as of 2026-05-03).

How to cite · data as of 2026-05-03

Grok-3 mini · benchmarks, pricing and providers. BenchGecko, data as of 2026-05-03. https://benchgecko.ai/model/grok-3-mini

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP