Grok 3 is the latest model from xAI. It's their flagship model that excels at enterprise use cases like data extraction, coding, and text summarization. Possesses deep domain knowledge in...
Tested on 6 benchmarks · BenchGecko score 67.9. Top scores: HELM — IFEval (88.4%), HELM — WildBench (84.9%), HELM — MMLU-Pro (78.8%).
DeepSeek V4 Flash scores 68.6 (101% as good) at $0.03/1M input · 99% cheaper
Multi-language code editing from Aider. Tests editing ability across Python, JavaScript, TypeScript, Java, C++, Go, Rust, and more.
Stanford HELM WildBench evaluation. Tests reasoning on challenging real-world tasks.
Stanford HELM evaluation of mathematical reasoning across diverse problem types.
- Typetext
- Context131K tokens (~66 books)
- ReleasedApr 2025
- LicenseProprietary
- Statuspreview
- Cost / Message~$0.021
Frequently Asked Questions
Key facts · as of 2026-05-03
- Grok 3 Beta by xAI. BenchGecko score 67.9, rank 47 of 312 scored models (normalized average of public benchmark scores).
- List price $3.00 input · $15.00 output per 1M tokens (as of 2026-05-03).
How to cite · data as of 2026-05-03
Grok 3 Beta · benchmarks, pricing and providers. BenchGecko, data as of 2026-05-03. https://benchgecko.ai/model/grok-3-beta
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP