Compare · ModelsLive · 3 picked · head to head

Claude 3.5 Sonnet vs DeepSeek V3 vs GPT-4 (older v0314)

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

DeepSeek V3 wins 9 of 19 shared benchmarks. Leads in math · reasoning.

Category leads
arena·Claude 3.5 Sonnetknowledge·Claude 3.5 Sonnetmath·DeepSeek V3coding·Claude 3.5 Sonnetgeneral·Claude 3.5 Sonnetlanguage·Claude 3.5 Sonnetreasoning·DeepSeek V3
Hype vs Reality
Claude 3.5 Sonnet
#190 by perf·no signal
QUIET
DeepSeek V3
#70 by perf·no signal
QUIET
GPT-4 (older v0314)
#81 by perf·no signal
QUIET
Best value
71.4x better value than GPT-4 (older v0314)
Claude 3.5 Sonnet
n/a
no price
DeepSeek V3
87.2 pts/$
$0.64/M
GPT-4 (older v0314)
1.2 pts/$
$45.00/M
Vendor risk
One or more vendors flagged
Anthropic logo
Anthropic
$965.0B·Tier 1
Medium risk
DeepSeek logo
DeepSeek
$3.4B·Tier 1
Higher risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Head to head
Claude 3.5 SonnetDeepSeek V3GPT-4 (older v0314)
Chatbot Arena Elo · Overall
Claude 3.5 Sonnet leads by +15.5
Claude 3.5 Sonnet
1374.0
DeepSeek V3
1358.4
GPT-4 (older v0314)
1285.8
GPQA diamond
DeepSeek V3 leads by +3.3
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
Claude 3.5 Sonnet
38.7
DeepSeek V3
42.0
GPT-4 (older v0314)
14.3
MMLU
DeepSeek V3 leads by +0.9
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Claude 3.5 Sonnet
82.0
DeepSeek V3
82.9
GPT-4 (older v0314)
81.9
OTIS Mock AIME 2024-2025
DeepSeek V3 leads by +9.3
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
Claude 3.5 Sonnet
6.4
DeepSeek V3
15.8
GPT-4 (older v0314)
0.5
Aider · Code Editing
Claude 3.5 Sonnet leads by +18.0
Claude 3.5 Sonnet
84.2
GPT-4 (older v0314)
66.2
Aider polyglot
Claude 3.5 Sonnet leads by +3.2
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
Claude 3.5 Sonnet
51.6
DeepSeek V3
48.4
Dtbench
Claude 3.5 Sonnet leads by +5.0
Claude 3.5 Sonnet
46.3
DeepSeek V3
41.3
FrontierMath-2025-02-28-Private
DeepSeek V3 leads by +1.2
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
Claude 3.5 Sonnet
1.8
DeepSeek V3
3.0
HELM · GPQA
Claude 3.5 Sonnet leads by +2.7
Claude 3.5 Sonnet
56.5
DeepSeek V3
53.8
HELM · IFEval
Claude 3.5 Sonnet leads by +2.4
Claude 3.5 Sonnet
85.6
DeepSeek V3
83.2
HELM · MMLU-Pro
Claude 3.5 Sonnet leads by +5.4
Claude 3.5 Sonnet
77.7
DeepSeek V3
72.3
HELM · Omni-MATH
DeepSeek V3 leads by +12.7
Claude 3.5 Sonnet
27.6
DeepSeek V3
40.3
HELM · WildBench
DeepSeek V3 leads by +3.9
Claude 3.5 Sonnet
79.2
DeepSeek V3
83.1
Lech Mazur Writing
Claude 3.5 Sonnet leads by +3.3
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
Claude 3.5 Sonnet
80.3
DeepSeek V3
77.0
MATH level 5
DeepSeek V3 leads by +13.2
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
Claude 3.5 Sonnet
51.7
DeepSeek V3
64.8
Metr Time Horizons
DeepSeek V3 leads by +7.2
Claude 3.5 Sonnet
40.1
DeepSeek V3
47.4
SimpleBench
Claude 3.5 Sonnet leads by +10.3
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
Claude 3.5 Sonnet
13.0
DeepSeek V3
2.7
WeirdML
DeepSeek V3 leads by +5.1
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
Claude 3.5 Sonnet
31.0
DeepSeek V3
36.1
Winogrande
GPT-4 (older v0314) leads by +4.6
WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs.
DeepSeek V3
70.4
GPT-4 (older v0314)
75.0
Full benchmark table
BenchmarkClaude 3.5 SonnetDeepSeek V3GPT-4 (older v0314)
Chatbot Arena Elo · Overall
1374.01358.41285.8
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
38.742.014.3
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
82.082.981.9
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
6.415.80.5
Aider · Code Editing
84.2—66.2
Aider polyglot
Aider Polyglot · measures how well AI models can edit code across multiple programming languages using the Aider coding assistant framework.
51.648.4—
Dtbench
46.341.3—
FrontierMath-2025-02-28-Private
FrontierMath (Feb 2025) · original research-level math problems created by mathematicians, testing capabilities at the boundary of current AI mathematical reasoning.
1.83.0—
HELM · GPQA
56.553.8—
HELM · IFEval
85.683.2—
HELM · MMLU-Pro
77.772.3—
HELM · Omni-MATH
27.640.3—
HELM · WildBench
79.283.1—
Lech Mazur Writing
Lech Mazur Writing · evaluates creative writing ability, assessing prose quality, narrative coherence, and stylistic sophistication.
80.377.0—
MATH level 5
MATH Level 5 · the hardest tier of the MATH benchmark, featuring competition-level problems from AMC, AIME, and Olympiad-style mathematics.
51.764.8—
Metr Time Horizons
40.147.4—
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
13.02.7—
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
31.036.1—
Winogrande
WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs.
—70.475.0
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Anthropic logoClaude 3.5 Sonnet————
DeepSeek logoDeepSeek V3$0.26$1.03164K tokens (~82 books)$4.50
OpenAI logoGPT-4 (older v0314)$30.00$60.008K tokens (~4 books)$375.00