Home/Models/Grok 4.20
xAI logo

Grok 4.20

by xAI · Released Mar 2026

Multimodal2M Context
56.5
avg score
Rank #108
Compare
Better than 65% of all models
Context
2.0M tokens (~1,000 books)
Input $/1M
$1.25
Output $/1M
$2.50
Type
multimodal
License
Proprietary
Benchmarks
15 tested
Data as of
About

Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

Tested on 15 benchmarks · BenchGecko score 56.5. Top scores: OTIS Mock AIME 2024-2025 (92.2%), ARC-AGI (89.5%), GPQA diamond (85.8%).

Looking for similar performance at lower cost?
Qwen2.5 Coder 7B Instruct scores 56.3 (100% as good) at $0.03/1M input · 98% cheaper
Capabilities
coding
54.8
#72 globally
reasoning
77.3
#24 globally
math
51.4
#116 globally
knowledge
45.3
#165 globally
general
35.4
#100 globally
Benchmark Scores
Compare All
Tested on 15 benchmarks · Ranked across 5 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
Terminal Bench

Complex terminal-based engineering tasks. Models must use command-line tools, navigate filesystems, and debug systems through shell interaction.

57.3·
WeirdML

Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.

52.3·
ARC-AGI

Abstraction and Reasoning Corpus. Tests fluid intelligence through novel visual pattern recognition puzzles. Core measure of general intelligence.

89.5·
ARC-AGI-2

ARC-AGI 2, harder sequel to ARC. More complex abstract reasoning patterns that test generalization ability beyond training data.

65.1·
OTIS Mock AIME 2024-2025

Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.

92.2·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Recently Happened
Grok 4.20 pricing dropped 58%
May 1, 2026
Grok 4.20 added
Apr 5, 2026
Links
Documentation
Community
BenchGecko API
grok-4-20
Specifications
  • Typemultimodal
  • Context2.0M tokens (~1,000 books)
  • ReleasedMar 2026
  • LicenseProprietary
  • StatusActive
  • Cost / Message~$0.005
Available On
xAI logoxAI$1.25
Share & Export
Tweet
Grok 4.20 is a proprietary multimodal AI model by xAI, released in March 2026. It has an average benchmark score of 56.5. Context window: 2M tokens.

Key facts · as of 2026-10-05

  • Grok 4.20 by xAI. BenchGecko score 56.5, rank 108 of 312 scored models (normalized average of public benchmark scores).
  • List price $1.25 input · $2.50 output per 1M tokens (as of 2026-10-05).
  • Sold by 4 providers (as of 2026-10-05): xAI $1.25 in / $2.50 out · xAI $1.25 in / $2.50 out · xAI $2.50 in / $5.00 out · xAI $2.50 in / $5.00 out. Every provider
  • Gecko Tests: Who Are You B (Knows who made it) · Tokenizer Tax B (44% more tokens outside English).

How to cite · data as of 2026-10-05

Grok 4.20 · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/grok-4-20

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP