Compare · ModelsLive · 3 picked · head to head

DeepSeek V4 Pro vs GPT-5.2-Codex vs MiniMax M3

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

DeepSeek V4 Pro wins 18 of 28 shared benchmarks. Leads in math · knowledge · speed.

Category leads
coding·MiniMax M3reasoning·GPT-5.2-Codexlanguage·GPT-5.2-Codexmath·DeepSeek V4 Proknowledge·DeepSeek V4 Prospeed·DeepSeek V4 Proarena·MiniMax M3general·DeepSeek V4 Pro
Hype vs Reality
DeepSeek V4 Pro
#89 by perf·#6 by attention
DESERVED
GPT-5.2-Codex
#17 by perf·#4 by attention
DESERVED
MiniMax M3
#124 by perf·#11 by attention
UNDERRATED
Best value
2.6x better value than MiniMax M3
DeepSeek V4 Pro
172.1 pts/$
$0.31/M
GPT-5.2-Codex
9.0 pts/$
$7.88/M
MiniMax M3
66.5 pts/$
$0.75/M
Vendor risk
One or more vendors flagged
DeepSeek logo
DeepSeek
$3.4B·Tier 1
Higher risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
minimax logo
MiniMax
$4.0B·Tier 1
Higher risk
Head to head
DeepSeek V4 ProGPT-5.2-CodexMiniMax M3
LiveBench · Agentic Coding
MiniMax M3 leads by +3.3
DeepSeek V4 Pro
56.7
GPT-5.2-Codex
51.7
MiniMax M3
60.0
LiveBench · Coding
GPT-5.2-Codex leads by +13.6
DeepSeek V4 Pro
70.0
GPT-5.2-Codex
83.6
MiniMax M3
68.2
LiveBench · Data Analysis
GPT-5.2-Codex leads by +2.0
DeepSeek V4 Pro
74.5
GPT-5.2-Codex
78.2
MiniMax M3
76.2
LiveBench · If
GPT-5.2-Codex leads by +4.1
DeepSeek V4 Pro
62.4
GPT-5.2-Codex
66.5
MiniMax M3
57.5
LiveBench · Language
DeepSeek V4 Pro leads by +1.3
DeepSeek V4 Pro
78.1
GPT-5.2-Codex
73.7
MiniMax M3
76.8
LiveBench · Mathematics
DeepSeek V4 Pro leads by +1.9
DeepSeek V4 Pro
90.7
GPT-5.2-Codex
88.8
MiniMax M3
77.0
LiveBench · Overall
GPT-5.2-Codex leads by +0.7
DeepSeek V4 Pro
73.6
GPT-5.2-Codex
74.3
MiniMax M3
70.0
LiveBench · Reasoning
DeepSeek V4 Pro leads by +5.0
DeepSeek V4 Pro
82.7
GPT-5.2-Codex
77.7
MiniMax M3
74.5
Artificial Analysis · Agentic Index
DeepSeek V4 Pro leads by +1.0
Artificial Analysis Agentic Index · a composite score measuring how well a model performs in agentic workflows · multi-step tool use, planning, error recovery, and autonomous task completion. Aggregates results from multiple agentic benchmarks including SWE-bench, tool-use tests, and planning evaluations. The canonical single-number metric for "how good is this model as an agent?"
DeepSeek V4 Pro
36.4
MiniMax M3
35.4
Artificial Analysis · Coding Index
DeepSeek V4 Pro leads by +0.8
Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads.
DeepSeek V4 Pro
59.4
MiniMax M3
58.6
Artificial Analysis · CritPt
DeepSeek V4 Pro leads by +14.3
DeepSeek V4 Pro
18.0
MiniMax M3
3.7
Artificial Analysis · GDPval
DeepSeek V4 Pro leads by +9.8
DeepSeek V4 Pro
47.1
MiniMax M3
37.3
Artificial Analysis · GPQA Diamond
MiniMax M3 leads by +0.1
DeepSeek V4 Pro
92.8
MiniMax M3
92.9
Artificial Analysis · Humanity's Last Exam
DeepSeek V4 Pro leads by +2.0
DeepSeek V4 Pro
41.0
MiniMax M3
39.0
Artificial Analysis · Long Context Reasoning
MiniMax M3 leads by +2.7
DeepSeek V4 Pro
80.3
MiniMax M3
83.0
Artificial Analysis · Quality Index
DeepSeek V4 Pro leads by +6.8
DeepSeek V4 Pro
36.0
MiniMax M3
29.2
Artificial Analysis · SciCode
DeepSeek V4 Pro leads by +3.9
DeepSeek V4 Pro
51.0
MiniMax M3
47.1
Chatbot Arena Elo · Coding
MiniMax M3 leads by +36.1
DeepSeek V4 Pro
1446.0
MiniMax M3
1482.0
Chatbot Arena Elo · Overall
DeepSeek V4 Pro leads by +17.7
DeepSeek V4 Pro
1457.8
MiniMax M3
1440.1
Chess Puzzles
DeepSeek V4 Pro leads by +6.3
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
DeepSeek V4 Pro
15.8
MiniMax M3
9.5
Dtbench
DeepSeek V4 Pro leads by +19.6
DeepSeek V4 Pro
84.5
MiniMax M3
64.9
Frontiercode
DeepSeek V4 Pro leads by +2.9
DeepSeek V4 Pro
17.6
MiniMax M3
14.7
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
DeepSeek V4 Pro
87.9
MiniMax M3
87.9
Lmca
DeepSeek V4 Pro leads by +8.8
DeepSeek V4 Pro
48.5
MiniMax M3
39.6
Mystery Game Puzzles
DeepSeek V4 Pro leads by +8.6
DeepSeek V4 Pro
8.6
MiniMax M3
0.0
OTIS Mock AIME 2024-2025
DeepSeek V4 Pro leads by +25.6
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
DeepSeek V4 Pro
96.7
MiniMax M3
71.1
Proofbench
MiniMax M3 leads by +2.0
DeepSeek V4 Pro
16.0
MiniMax M3
18.0
Surface Evolver Bench
MiniMax M3 leads by +15.0
DeepSeek V4 Pro
40.0
MiniMax M3
55.0
Full benchmark table
BenchmarkDeepSeek V4 ProGPT-5.2-CodexMiniMax M3
LiveBench · Agentic Coding
56.751.760.0
LiveBench · Coding
70.083.668.2
LiveBench · Data Analysis
74.578.276.2
LiveBench · If
62.466.557.5
LiveBench · Language
78.173.776.8
LiveBench · Mathematics
90.788.877.0
LiveBench · Overall
73.674.370.0
LiveBench · Reasoning
82.777.774.5
Artificial Analysis · Agentic Index
Artificial Analysis Agentic Index · a composite score measuring how well a model performs in agentic workflows · multi-step tool use, planning, error recovery, and autonomous task completion. Aggregates results from multiple agentic benchmarks including SWE-bench, tool-use tests, and planning evaluations. The canonical single-number metric for "how good is this model as an agent?"
36.4—35.4
Artificial Analysis · Coding Index
Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads.
59.4—58.6
Artificial Analysis · CritPt
18.0—3.7
Artificial Analysis · GDPval
47.1—37.3
Artificial Analysis · GPQA Diamond
92.8—92.9
Artificial Analysis · Humanity's Last Exam
41.0—39.0
Artificial Analysis · Long Context Reasoning
80.3—83.0
Artificial Analysis · Quality Index
36.0—29.2
Artificial Analysis · SciCode
51.0—47.1
Chatbot Arena Elo · Coding
1446.0—1482.0
Chatbot Arena Elo · Overall
1457.8—1440.1
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
15.8—9.5
Dtbench
84.5—64.9
Frontiercode
17.6—14.7
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
87.9—87.9
Lmca
48.5—39.6
Mystery Game Puzzles
8.6—0.0
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
96.7—71.1
Proofbench
16.0—18.0
Surface Evolver Bench
40.0—55.0
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
DeepSeek logoDeepSeek V4 Pro$0.21$0.421.0M tokens (~524 books)$2.61
OpenAI logoGPT-5.2-Codex$1.75$14.00400K tokens (~200 books)$48.13
minimax logoMiniMax M3$0.30$1.201.0M tokens (~524 books)$5.25