The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.
Tested on 18 benchmarks · BenchGecko score 40.9. Top scores: HellaSwag (93.7%), GSM8K (90.0%), TriviaQA (84.8%).
Qwen3 30B A3B Instruct 2507 scores 40.1 (98% as good) at $0.05/1M input · 100% cheaper
Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.
BIG-Bench Hard. 23 challenging tasks from BIG-Bench where prior language models fell below average human performance.
Deceptively simple questions that humans find easy but AI models often get wrong. Tests common sense and reasoning gaps.
Grade school math word problems. 8,500 problems testing multi-step arithmetic reasoning. A foundational math benchmark.
Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
- Typemultimodal
- Context128K tokens (~64 books)
- ReleasedApr 2024
- LicenseProprietary
- StatusActive
- Cost / Message~$0.050
Frequently Asked Questions
Key facts · as of 2026-10-05
- GPT-4 Turbo by OpenAI. BenchGecko score 40.9, rank 194 of 312 scored models (normalized average of public benchmark scores).
- List price $10.00 input · $30.00 output per 1M tokens (as of 2026-10-05).
- Sold by 1 provider (as of 2026-10-05): OpenAI $10.00 in / $30.00 out.
How to cite · data as of 2026-10-05
GPT-4 Turbo · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/gpt-4-turbo
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP