Home/Models/DeepSeek V3
DeepSeek logo

DeepSeek V3

by DeepSeek · Released Dec 2024

Open Source
55.1
avg score
Rank #118
Compare
Better than 62% of all models
Context
164K tokens (~82 books)
Input $/1M
$0.26
Output $/1M
$1.03
Type
text
License
Open Source
Benchmarks
25 tested
Data as of
About

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions. Pre-trained on nearly 15 trillion tokens, the reported evaluations...

Tested on 25 benchmarks · BenchGecko score 55.1. Top scores: Chatbot Arena Elo — Overall (1358.4%), ARC AI2 (93.7%), HellaSwag (85.2%).

Looking for similar performance at lower cost?
Qwen3 235B A22B Thinking 2507 scores 55.3 (100% as good) at $0.23/1M input · 11% cheaper
Capabilities
coding
42.2
#129 globally
reasoning
56.4
#68 globally
math
31.0
#185 globally
knowledge
70.9
#17 globally
general
35.6
#95 globally
language
83.2
#47 globally
Benchmark Scores
Compare All
Tested on 25 benchmarks · Ranked across 7 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
Aider polyglot

Multi-language code editing from Aider. Tests editing ability across Python, JavaScript, TypeScript, Java, C++, Go, Rust, and more.

48.4·
WeirdML

Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.

36.1·
BBH

BIG-Bench Hard. 23 challenging tasks from BIG-Bench where prior language models fell below average human performance.

83.3·
HELM — WildBench

Stanford HELM WildBench evaluation. Tests reasoning on challenging real-world tasks.

83.1·
SimpleBench

Deceptively simple questions that humans find easy but AI models often get wrong. Tests common sense and reasoning gaps.

2.7·
MATH level 5

Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.

64.8·
HELM — Omni-MATH

Stanford HELM evaluation of mathematical reasoning across diverse problem types.

40.3·
OTIS Mock AIME 2024-2025

Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.

15.8·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Links
Documentation
Community
BenchGecko API
deepseek-chat
Specifications
  • Typetext
  • Context164K tokens (~82 books)
  • ReleasedDec 2024
  • LicenseOpen Source
  • StatusActive
  • Cost / Message~$0.002
Available On
DeepSeek logoDeepSeek$0.26
Share & Export
Tweet
DeepSeek V3 is an open-source text AI model by DeepSeek, released in December 2024. It has an average benchmark score of 55.1. Context window: 164K tokens.

Key facts · as of 2026-10-05

  • DeepSeek V3 by DeepSeek. BenchGecko score 55.1, rank 117 of 312 scored models (normalized average of public benchmark scores).
  • List price $0.26 input · $1.03 output per 1M tokens (as of 2026-10-05).
  • Sold by 2 providers (as of 2026-10-05): StreamLake $0.26 in / $1.03 out · DeepInfra (fp4) $0.32 in / $0.89 out. Every provider

How to cite · data as of 2026-10-05

DeepSeek V3 · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/deepseek-chat

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP