Compare · ModelsLive · 2 picked · head to head
Gemini 3.8 Flash vs GPT-5.4 Pro
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Gemini 3.8 Flash wins on 4/6 benchmarks
Gemini 3.8 Flash wins 4 of 6 shared benchmarks. Leads in knowledge.
Category leads
knowledge·Gemini 3.8 Flashmath·GPT-5.4 Pro
Hype vs Reality
Attention vs performance
Gemini 3.8 Flash
#63 by perf·#5 by attention
GPT-5.4 Pro
#37 by perf·#4 by attention
Best value
Gemini 3.8 Flash
41.8x better value than GPT-5.4 Pro
Gemini 3.8 Flash
25.5 pts/$
$2.25/M
GPT-5.4 Pro
0.6 pts/$
$105.00/M
Vendor risk
Who is behind the model
Google DeepMind
$4.20T·Tier 1
OpenAI
$840.0B·Tier 1
Head to head
6 benchmarks · 2 models
Gemini 3.8 FlashGPT-5.4 Pro
Chess Puzzles
Gemini 3.8 Flash leads by +2.5
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
Gemini 3.8 Flash
59.0
GPT-5.4 Pro
56.4
FrontierMath-Tier-4-v2-Private
GPT-5.4 Pro leads by +36.6
Gemini 3.8 Flash
21.9
GPT-5.4 Pro
58.5
FrontierMath-Tiers-1-3-v2-Private
GPT-5.4 Pro leads by +14.0
Gemini 3.8 Flash
68.4
GPT-5.4 Pro
82.5
GPQA diamond
Gemini 3.8 Flash leads by +1.1
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Gemini 3.8 Flash
93.9
GPT-5.4 Pro
92.8
HLE
Gemini 3.8 Flash leads by +0.2
HLE (Humanity's Last Exam) · a reasoning benchmark designed to be the hardest public evaluation of AI. Questions span mathematics, physics, philosophy, and logic · curated to be at or beyond the frontier of human expert capability. Tested with and without tool augmentation. Claude Opus 4.7 scores 46.9% without tools and 54.7% with tools · making it one of the few benchmarks where the top score is below 60%.
Gemini 3.8 Flash
41.7
GPT-5.4 Pro
41.5
SimpleQA Verified
Gemini 3.8 Flash leads by +23.4
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
Gemini 3.8 Flash
69.7
GPT-5.4 Pro
46.3
Full benchmark table
| Benchmark | Gemini 3.8 Flash | GPT-5.4 Pro |
|---|---|---|
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 59.0 | 56.4 |
FrontierMath-Tier-4-v2-Private | 21.9 | 58.5 |
FrontierMath-Tiers-1-3-v2-Private | 68.4 | 82.5 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 93.9 | 92.8 |
HLE HLE (Humanity's Last Exam) · a reasoning benchmark designed to be the hardest public evaluation of AI. Questions span mathematics, physics, philosophy, and logic · curated to be at or beyond the frontier of human expert capability. Tested with and without tool augmentation. Claude Opus 4.7 scores 46.9% without tools and 54.7% with tools · making it one of the few benchmarks where the top score is below 60%. | 41.7 | 41.5 |
SimpleQA Verified SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information. | 69.7 | 46.3 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.75 | $3.75 | 1.0M tokens (~524 books) | $15.00 | |
| $30.00 | $180.00 | 1.1M tokens (~525 books) | $675.00 |