Compare · ModelsLive · 2 picked · head to head
gpt-oss-120b vs Qwen3.6 Flash
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Qwen3.6 Flash wins on 13/16 benchmarks
Qwen3.6 Flash wins 13 of 16 shared benchmarks. Leads in knowledge · general · coding.
Category leads
knowledge·Qwen3.6 Flashgeneral·Qwen3.6 Flashcoding·Qwen3.6 Flashreasoning·Qwen3.6 Flashlanguage·gpt-oss-120bmath·Qwen3.6 Flash
Hype vs Reality
Attention vs performance
gpt-oss-120b
#127 by perf·no signal
Qwen3.6 Flash
#168 by perf·#2 by attention
Best value
gpt-oss-120b
7.1x better value than Qwen3.6 Flash
gpt-oss-120b
475.4 pts/$
$0.10/M
Qwen3.6 Flash
67.4 pts/$
$0.66/M
Vendor risk
Who is behind the model
OpenAI
$840.0B·Tier 1
Alibaba (Qwen)
$293.0B·Tier 1
Head to head
16 benchmarks · 2 models
gpt-oss-120bQwen3.6 Flash
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
gpt-oss-120b
15.8
Qwen3.6 Flash
15.8
Dtbench
Qwen3.6 Flash leads by +1.3
gpt-oss-120b
60.5
Qwen3.6 Flash
61.8
GPQA diamond
Qwen3.6 Flash leads by +10.1
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
gpt-oss-120b
67.7
Qwen3.6 Flash
77.8
LiveBench · Agentic Coding
Qwen3.6 Flash leads by +30.0
gpt-oss-120b
16.7
Qwen3.6 Flash
46.7
LiveBench · Coding
Qwen3.6 Flash leads by +4.7
gpt-oss-120b
60.2
Qwen3.6 Flash
64.9
LiveBench · Data Analysis
Qwen3.6 Flash leads by +20.0
gpt-oss-120b
38.8
Qwen3.6 Flash
58.8
LiveBench · If
gpt-oss-120b leads by +3.1
gpt-oss-120b
50.3
Qwen3.6 Flash
47.2
LiveBench · Language
Qwen3.6 Flash leads by +14.6
gpt-oss-120b
48.6
Qwen3.6 Flash
63.1
LiveBench · Mathematics
Qwen3.6 Flash leads by +10.0
gpt-oss-120b
68.9
Qwen3.6 Flash
78.9
LiveBench · Overall
Qwen3.6 Flash leads by +14.3
gpt-oss-120b
46.1
Qwen3.6 Flash
60.4
LiveBench · Reasoning
Qwen3.6 Flash leads by +23.7
gpt-oss-120b
39.2
Qwen3.6 Flash
62.9
Lmca
Qwen3.6 Flash leads by +10.4
gpt-oss-120b
26.1
Qwen3.6 Flash
36.5
Mystery Game Puzzles
Qwen3.6 Flash leads by +9.7
gpt-oss-120b
0.0
Qwen3.6 Flash
9.7
OTIS Mock AIME 2024-2025
gpt-oss-120b leads by +4.4
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
gpt-oss-120b
88.9
Qwen3.6 Flash
84.4
SimpleBench
Qwen3.6 Flash leads by +15.7
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
gpt-oss-120b
6.5
Qwen3.6 Flash
22.2
SimpleQA Verified
Qwen3.6 Flash leads by +2.0
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
gpt-oss-120b
13.9
Qwen3.6 Flash
15.9
Full benchmark table
| Benchmark | gpt-oss-120b | Qwen3.6 Flash |
|---|---|---|
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 15.8 | 15.8 |
Dtbench | 60.5 | 61.8 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 67.7 | 77.8 |
LiveBench · Agentic Coding | 16.7 | 46.7 |
LiveBench · Coding | 60.2 | 64.9 |
LiveBench · Data Analysis | 38.8 | 58.8 |
LiveBench · If | 50.3 | 47.2 |
LiveBench · Language | 48.6 | 63.1 |
LiveBench · Mathematics | 68.9 | 78.9 |
LiveBench · Overall | 46.1 | 60.4 |
LiveBench · Reasoning | 39.2 | 62.9 |
Lmca | 26.1 | 36.5 |
Mystery Game Puzzles | 0.0 | 9.7 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 88.9 | 84.4 |
SimpleBench SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking. | 6.5 | 22.2 |
SimpleQA Verified SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information. | 13.9 | 15.9 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.04 | $0.17 | 131K tokens (~66 books) | $0.70 | |
| $0.19 | $1.13 | 1.0M tokens (~500 books) | $4.22 |