Compare · ModelsLive · 2 picked · head to head
Qwen3.5 397B A17B vs Qwen3 Next 80B A3B Instruct
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Qwen3.5 397B A17B wins on 17/18 benchmarks
Qwen3.5 397B A17B wins 17 of 18 shared benchmarks. Leads in speed · arena · math.
Category leads
speed·Qwen3.5 397B A17Barena·Qwen3.5 397B A17Bmath·Qwen3.5 397B A17Bknowledge·Qwen3.5 397B A17Blanguage·Qwen3.5 397B A17Bcoding·Qwen3.5 397B A17B
Hype vs Reality
Attention vs performance
Qwen3.5 397B A17B
#48 by perf·#2 by attention
Qwen3 Next 80B A3B Instruct
#83 by perf·#2 by attention
Best value
Qwen3 Next 80B A3B Instruct
3.1x better value than Qwen3.5 397B A17B
Qwen3.5 397B A17B
29.6 pts/$
$2.02/M
Qwen3 Next 80B A3B Instruct
91.4 pts/$
$0.60/M
Vendor risk
Who is behind the model
Alibaba (Qwen)
$293.0B·Tier 1
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
18 benchmarks · 2 models
Qwen3.5 397B A17BQwen3 Next 80B A3B Instruct
Artificial Analysis · Agentic Index
Qwen3.5 397B A17B leads by +5.7
Artificial Analysis Agentic Index · a composite score measuring how well a model performs in agentic workflows · multi-step tool use, planning, error recovery, and autonomous task completion. Aggregates results from multiple agentic benchmarks including SWE-bench, tool-use tests, and planning evaluations. The canonical single-number metric for "how good is this model as an agent?"
Qwen3.5 397B A17B
19.9
Qwen3 Next 80B A3B Instruct
14.2
Artificial Analysis · Coding Index
Qwen3.5 397B A17B leads by +32.9
Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads.
Qwen3.5 397B A17B
48.2
Qwen3 Next 80B A3B Instruct
15.3
Artificial Analysis · CritPt
Qwen3.5 397B A17B leads by +0.9
Qwen3.5 397B A17B
0.9
Qwen3 Next 80B A3B Instruct
0.0
Artificial Analysis · GDPval
Qwen3.5 397B A17B leads by +14.0
Qwen3.5 397B A17B
14.0
Qwen3 Next 80B A3B Instruct
0.0
Artificial Analysis · GPQA Diamond
Qwen3.5 397B A17B leads by +10.2
Qwen3.5 397B A17B
86.1
Qwen3 Next 80B A3B Instruct
75.9
Artificial Analysis · Humanity's Last Exam
Qwen3.5 397B A17B leads by +7.2
Qwen3.5 397B A17B
19.8
Qwen3 Next 80B A3B Instruct
12.6
Artificial Analysis · IFBench
Qwen3 Next 80B A3B Instruct leads by +9.1
Qwen3.5 397B A17B
51.6
Qwen3 Next 80B A3B Instruct
60.7
Artificial Analysis · Long Context Reasoning
Qwen3.5 397B A17B leads by +0.6
Qwen3.5 397B A17B
64.3
Qwen3 Next 80B A3B Instruct
63.7
Artificial Analysis · Quality Index
Qwen3.5 397B A17B leads by +10.2
Qwen3.5 397B A17B
21.4
Qwen3 Next 80B A3B Instruct
11.2
Artificial Analysis · tau2-Bench Telecom
Qwen3.5 397B A17B leads by +42.4
Qwen3.5 397B A17B
83.9
Qwen3 Next 80B A3B Instruct
41.5
Artificial Analysis · Terminal-Bench Hard
Qwen3.5 397B A17B leads by +25.8
Qwen3.5 397B A17B
35.6
Qwen3 Next 80B A3B Instruct
9.8
Chatbot Arena Elo · Overall
Qwen3.5 397B A17B leads by +42.8
Qwen3.5 397B A17B
1441.8
Qwen3 Next 80B A3B Instruct
1399.0
OpenCompass · AIME2025
Qwen3.5 397B A17B leads by +23.1
Qwen3.5 397B A17B
92.3
Qwen3 Next 80B A3B Instruct
69.2
OpenCompass · GPQA-Diamond
Qwen3.5 397B A17B leads by +14.3
Qwen3.5 397B A17B
88.4
Qwen3 Next 80B A3B Instruct
74.1
OpenCompass · HLE
Qwen3.5 397B A17B leads by +19.5
Qwen3.5 397B A17B
27.5
Qwen3 Next 80B A3B Instruct
8.0
OpenCompass · IFEval
Qwen3.5 397B A17B leads by +3.9
Qwen3.5 397B A17B
91.5
Qwen3 Next 80B A3B Instruct
87.6
OpenCompass · LiveCodeBenchV6
Qwen3.5 397B A17B leads by +28.2
Qwen3.5 397B A17B
83.0
Qwen3 Next 80B A3B Instruct
54.8
OpenCompass · MMLU-Pro
Qwen3.5 397B A17B leads by +6.3
Qwen3.5 397B A17B
87.6
Qwen3 Next 80B A3B Instruct
81.3
Full benchmark table
| Benchmark | Qwen3.5 397B A17B | Qwen3 Next 80B A3B Instruct |
|---|---|---|
Artificial Analysis · Agentic Index Artificial Analysis Agentic Index · a composite score measuring how well a model performs in agentic workflows · multi-step tool use, planning, error recovery, and autonomous task completion. Aggregates results from multiple agentic benchmarks including SWE-bench, tool-use tests, and planning evaluations. The canonical single-number metric for "how good is this model as an agent?" | 19.9 | 14.2 |
Artificial Analysis · Coding Index Artificial Analysis Coding Index · a composite score that aggregates performance across multiple coding benchmarks into a single index. Tracks code generation quality, debugging ability, multi-language competence, and real-world software engineering tasks. Used by Artificial Analysis to rank model coding capability in a normalized, comparable format. Useful for developers choosing between models for coding-heavy workloads. | 48.2 | 15.3 |
Artificial Analysis · CritPt | 0.9 | 0.0 |
Artificial Analysis · GDPval | 14.0 | 0.0 |
Artificial Analysis · GPQA Diamond | 86.1 | 75.9 |
Artificial Analysis · Humanity's Last Exam | 19.8 | 12.6 |
Artificial Analysis · IFBench | 51.6 | 60.7 |
Artificial Analysis · Long Context Reasoning | 64.3 | 63.7 |
Artificial Analysis · Quality Index | 21.4 | 11.2 |
Artificial Analysis · tau2-Bench Telecom | 83.9 | 41.5 |
Artificial Analysis · Terminal-Bench Hard | 35.6 | 9.8 |
Chatbot Arena Elo · Overall | 1441.8 | 1399.0 |
OpenCompass · AIME2025 | 92.3 | 69.2 |
OpenCompass · GPQA-Diamond | 88.4 | 74.1 |
OpenCompass · HLE | 27.5 | 8.0 |
OpenCompass · IFEval | 91.5 | 87.6 |
OpenCompass · LiveCodeBenchV6 | 83.0 | 54.8 |
OpenCompass · MMLU-Pro | 87.6 | 81.3 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.55 | $3.50 | 262K tokens (~131 books) | $12.88 | |
| $0.09 | $1.10 | 262K tokens (~131 books) | $3.43 |