Gemini 1.5 Flash (May 2024)
by Google DeepMind · Released Jan 2024
Tested on 18 benchmarks · BenchGecko score 41.7. Top scores: Chatbot Arena Elo — Overall (1286.5%), HELM — IFEval (83.1%), GSM8K (82.4%).
Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.
Stanford HELM WildBench evaluation. Tests reasoning on challenging real-world tasks.
Grade school math word problems. 8,500 problems testing multi-step arithmetic reasoning. A foundational math benchmark.
Stanford HELM evaluation of mathematical reasoning across diverse problem types.
Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.
- Typetext
- ContextN/A
- ReleasedJan 2024
- LicenseProprietary
- Statusbenchmark-only
Frequently Asked Questions
Key facts · as of 2026-04-09
- Gemini 1.5 Flash (May 2024) by Google DeepMind. BenchGecko score 41.7, rank 190 of 312 scored models (normalized average of public benchmark scores).
- List price n/a input · n/a output per 1M tokens (as of 2026-04-09).
How to cite · data as of 2026-04-09
Gemini 1.5 Flash (May 2024) · benchmarks, pricing and providers. BenchGecko, data as of 2026-04-09. https://benchgecko.ai/model/gemini-1-5-flash-may-2024
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP