GPT-4-0314 is the first version of GPT-4 released, with a context length of 8,192 tokens, and was supported until June 14. Training data: up to Sep 2021.
Tested on 7 benchmarks · BenchGecko score 63.8. Top scores: Chatbot Arena Elo — Overall (1285.8%), GSM8K (92.0%), MMLU (81.9%).
Qwen3 30B A3B Thinking 2507 scores 63.5 (100% as good) at $0.20/1M input · 99% cheaper
Code editing benchmark from the Aider project. Measures ability to apply targeted code changes while maintaining correctness and style.
Grade school math word problems. 8,500 problems testing multi-step arithmetic reasoning. A foundational math benchmark.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
Massive Multitask Language Understanding. 57 subjects from STEM, humanities, and social sciences. The most widely-cited knowledge benchmark.
Commonsense coreference resolution. Tests understanding of pronoun references in ambiguous sentences.
Graduate-level science questions written by PhD experts. Diamond subset contains questions where experts disagree, testing deep understanding.
- Typetext
- Context8K tokens (~4 books)
- ReleasedMay 2023
- LicenseProprietary
- StatusActive
- Cost / Message~$0.120
Frequently Asked Questions
Key facts · as of 2026-05-03
- GPT-4 (older v0314) by OpenAI. BenchGecko score 63.8, rank 65 of 312 scored models (normalized average of public benchmark scores).
- List price $30.00 input · $60.00 output per 1M tokens (as of 2026-05-03).
How to cite · data as of 2026-05-03
GPT-4 (older v0314) · benchmarks, pricing and providers. BenchGecko, data as of 2026-05-03. https://benchgecko.ai/model/gpt-4-0314
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP