DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations...
Tested on 25 benchmarks · BenchGecko score 55.1. Top scores: Chatbot Arena Elo — Overall (1358.4%), ARC AI2 (93.7%), HellaSwag (85.2%).
Qwen3 235B A22B Thinking 2507 scores 55.3 (100% as good) at $0.23/1M input · 11% cheaper
Multi-language code editing from Aider. Tests editing ability across Python, JavaScript, TypeScript, Java, C++, Go, Rust, and more.
Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.
BIG-Bench Hard. 23 challenging tasks from BIG-Bench where prior language models fell below average human performance.
Stanford HELM WildBench evaluation. Tests reasoning on challenging real-world tasks.
Deceptively simple questions that humans find easy but AI models often get wrong. Tests common sense and reasoning gaps.
Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.
Stanford HELM evaluation of mathematical reasoning across diverse problem types.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
- Typetext
- Context164K tokens (~82 books)
- ReleasedDec 2024
- LicenseOpen Source
- StatusActive
- Cost / Message~$0.002
Frequently Asked Questions
Key facts · as of 2026-10-05
- DeepSeek V3 by DeepSeek. BenchGecko score 55.1, rank 117 of 312 scored models (normalized average of public benchmark scores).
- List price $0.26 input · $1.03 output per 1M tokens (as of 2026-10-05).
- Sold by 2 providers (as of 2026-10-05): StreamLake $0.26 in / $1.03 out · DeepInfra (fp4) $0.32 in / $0.89 out. Every provider
How to cite · data as of 2026-10-05
DeepSeek V3 · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/deepseek-chat
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP