Compare · ModelsLive · 2 picked · head to head
DeepSeek V4 Pro vs MiniMax M3
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
DeepSeek V4 Pro wins on 21/28 benchmarks
DeepSeek V4 Pro wins 21 of 28 shared benchmarks. Leads in speed · knowledge · general.
Category leads
speed·DeepSeek V4 Proarena·MiniMax M3knowledge·DeepSeek V4 Progeneral·DeepSeek V4 Procoding·DeepSeek V4 Proreasoning·MiniMax M3language·DeepSeek V4 Promath·DeepSeek V4 Pro
Hype vs Reality
Attention vs performance
DeepSeek V4 Pro
#89 by perf·#6 by attention
MiniMax M3
#124 by perf·#11 by attention
Best value
DeepSeek V4 Pro
2.6x better value than MiniMax M3
DeepSeek V4 Pro
172.1 pts/$
$0.31/M
MiniMax M3
66.5 pts/$
$0.75/M
Vendor risk
Mixed exposure
One or more vendors flagged
DeepSeek
$3.4B·Tier 1
MiniMax
$4.0B·Tier 1
Head to head
28 benchmarks · 2 models
DeepSeek V4 ProMiniMax M3
Artificial Analysis · Agentic Index
DeepSeek V4 Pro leads by +1.0
Artificial Analysis Agentic Index · a composite score measuring how well a model performs in agentic workflows · multi-step tool use, planning, error recovery, and autonomous task completion. Aggregates results from multiple agentic benchmarks including SWE-bench, tool-use tests, and planning evaluations. The canonical single-number metric for "how good is this model as an agent?"
DeepSeek V4 Pro
36.4
MiniMax M3
35.4
Artificial Analysis · Coding Index
DeepSeek V4 Pro leads by +0.8
Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads.
DeepSeek V4 Pro
59.4
MiniMax M3
58.6
Artificial Analysis · CritPt
DeepSeek V4 Pro leads by +14.3
DeepSeek V4 Pro
18.0
MiniMax M3
3.7
Artificial Analysis · GDPval
DeepSeek V4 Pro leads by +9.8
DeepSeek V4 Pro
47.1
MiniMax M3
37.3
Artificial Analysis · GPQA Diamond
MiniMax M3 leads by +0.1
DeepSeek V4 Pro
92.8
MiniMax M3
92.9
Artificial Analysis · Humanity's Last Exam
DeepSeek V4 Pro leads by +2.0
DeepSeek V4 Pro
41.0
MiniMax M3
39.0
Artificial Analysis · Long Context Reasoning
MiniMax M3 leads by +2.7
DeepSeek V4 Pro
80.3
MiniMax M3
83.0
Artificial Analysis · Quality Index
DeepSeek V4 Pro leads by +6.8
DeepSeek V4 Pro
36.0
MiniMax M3
29.2
Artificial Analysis · SciCode
DeepSeek V4 Pro leads by +3.9
DeepSeek V4 Pro
51.0
MiniMax M3
47.1
Chatbot Arena Elo · Coding
MiniMax M3 leads by +36.1
DeepSeek V4 Pro
1446.0
MiniMax M3
1482.0
Chatbot Arena Elo · Overall
DeepSeek V4 Pro leads by +17.7
DeepSeek V4 Pro
1457.8
MiniMax M3
1440.1
Chess Puzzles
DeepSeek V4 Pro leads by +6.3
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
DeepSeek V4 Pro
15.8
MiniMax M3
9.5
Dtbench
DeepSeek V4 Pro leads by +19.6
DeepSeek V4 Pro
84.5
MiniMax M3
64.9
Frontiercode
DeepSeek V4 Pro leads by +2.9
DeepSeek V4 Pro
17.6
MiniMax M3
14.7
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
DeepSeek V4 Pro
87.9
MiniMax M3
87.9
LiveBench · Agentic Coding
MiniMax M3 leads by +3.3
DeepSeek V4 Pro
56.7
MiniMax M3
60.0
LiveBench · Coding
DeepSeek V4 Pro leads by +1.8
DeepSeek V4 Pro
70.0
MiniMax M3
68.2
LiveBench · Data Analysis
MiniMax M3 leads by +1.6
DeepSeek V4 Pro
74.5
MiniMax M3
76.2
LiveBench · If
DeepSeek V4 Pro leads by +4.8
DeepSeek V4 Pro
62.4
MiniMax M3
57.5
LiveBench · Language
DeepSeek V4 Pro leads by +1.3
DeepSeek V4 Pro
78.1
MiniMax M3
76.8
LiveBench · Mathematics
DeepSeek V4 Pro leads by +13.7
DeepSeek V4 Pro
90.7
MiniMax M3
77.0
LiveBench · Overall
DeepSeek V4 Pro leads by +3.6
DeepSeek V4 Pro
73.6
MiniMax M3
70.0
LiveBench · Reasoning
DeepSeek V4 Pro leads by +8.2
DeepSeek V4 Pro
82.7
MiniMax M3
74.5
Lmca
DeepSeek V4 Pro leads by +8.8
DeepSeek V4 Pro
48.5
MiniMax M3
39.6
Mystery Game Puzzles
DeepSeek V4 Pro leads by +8.6
DeepSeek V4 Pro
8.6
MiniMax M3
0.0
OTIS Mock AIME 2024-2025
DeepSeek V4 Pro leads by +25.6
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
DeepSeek V4 Pro
96.7
MiniMax M3
71.1
Proofbench
MiniMax M3 leads by +2.0
DeepSeek V4 Pro
16.0
MiniMax M3
18.0
Surface Evolver Bench
MiniMax M3 leads by +15.0
DeepSeek V4 Pro
40.0
MiniMax M3
55.0
Full benchmark table
| Benchmark | DeepSeek V4 Pro | MiniMax M3 |
|---|---|---|
Artificial Analysis · Agentic Index Artificial Analysis Agentic Index · a composite score measuring how well a model performs in agentic workflows · multi-step tool use, planning, error recovery, and autonomous task completion. Aggregates results from multiple agentic benchmarks including SWE-bench, tool-use tests, and planning evaluations. The canonical single-number metric for "how good is this model as an agent?" | 36.4 | 35.4 |
Artificial Analysis · Coding Index Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads. | 59.4 | 58.6 |
Artificial Analysis · CritPt | 18.0 | 3.7 |
Artificial Analysis · GDPval | 47.1 | 37.3 |
Artificial Analysis · GPQA Diamond | 92.8 | 92.9 |
Artificial Analysis · Humanity's Last Exam | 41.0 | 39.0 |
Artificial Analysis · Long Context Reasoning | 80.3 | 83.0 |
Artificial Analysis · Quality Index | 36.0 | 29.2 |
Artificial Analysis · SciCode | 51.0 | 47.1 |
Chatbot Arena Elo · Coding | 1446.0 | 1482.0 |
Chatbot Arena Elo · Overall | 1457.8 | 1440.1 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 15.8 | 9.5 |
Dtbench | 84.5 | 64.9 |
Frontiercode | 17.6 | 14.7 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 87.9 | 87.9 |
LiveBench · Agentic Coding | 56.7 | 60.0 |
LiveBench · Coding | 70.0 | 68.2 |
LiveBench · Data Analysis | 74.5 | 76.2 |
LiveBench · If | 62.4 | 57.5 |
LiveBench · Language | 78.1 | 76.8 |
LiveBench · Mathematics | 90.7 | 77.0 |
LiveBench · Overall | 73.6 | 70.0 |
LiveBench · Reasoning | 82.7 | 74.5 |
Lmca | 48.5 | 39.6 |
Mystery Game Puzzles | 8.6 | 0.0 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 96.7 | 71.1 |
Proofbench | 16.0 | 18.0 |
Surface Evolver Bench | 40.0 | 55.0 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.21 | $0.42 | 1.0M tokens (~524 books) | $2.61 | |
| $0.30 | $1.20 | 1.0M tokens (~524 books) | $5.25 |
People also compared