Compare · ModelsLive · 2 picked · head to head

gpt-oss-120b vs Qwen3.6 Flash

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Qwen3.6 Flash wins 13 of 16 shared benchmarks. Leads in knowledge · general · coding.

Category leads
knowledge·Qwen3.6 Flashgeneral·Qwen3.6 Flashcoding·Qwen3.6 Flashreasoning·Qwen3.6 Flashlanguage·gpt-oss-120bmath·Qwen3.6 Flash
Hype vs Reality
gpt-oss-120b
#127 by perf·no signal
QUIET
Qwen3.6 Flash
#168 by perf·#2 by attention
OVERHYPED
Best value
7.1x better value than Qwen3.6 Flash
gpt-oss-120b
475.4 pts/$
$0.10/M
Qwen3.6 Flash
67.4 pts/$
$0.66/M
Vendor risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Alibaba Qwen logo
Alibaba (Qwen)
$293.0B·Tier 1
Low risk
Head to head
gpt-oss-120bQwen3.6 Flash
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
gpt-oss-120b
15.8
Qwen3.6 Flash
15.8
Dtbench
Qwen3.6 Flash leads by +1.3
gpt-oss-120b
60.5
Qwen3.6 Flash
61.8
GPQA diamond
Qwen3.6 Flash leads by +10.1
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
gpt-oss-120b
67.7
Qwen3.6 Flash
77.8
LiveBench · Agentic Coding
Qwen3.6 Flash leads by +30.0
gpt-oss-120b
16.7
Qwen3.6 Flash
46.7
LiveBench · Coding
Qwen3.6 Flash leads by +4.7
gpt-oss-120b
60.2
Qwen3.6 Flash
64.9
LiveBench · Data Analysis
Qwen3.6 Flash leads by +20.0
gpt-oss-120b
38.8
Qwen3.6 Flash
58.8
LiveBench · If
gpt-oss-120b leads by +3.1
gpt-oss-120b
50.3
Qwen3.6 Flash
47.2
LiveBench · Language
Qwen3.6 Flash leads by +14.6
gpt-oss-120b
48.6
Qwen3.6 Flash
63.1
LiveBench · Mathematics
Qwen3.6 Flash leads by +10.0
gpt-oss-120b
68.9
Qwen3.6 Flash
78.9
LiveBench · Overall
Qwen3.6 Flash leads by +14.3
gpt-oss-120b
46.1
Qwen3.6 Flash
60.4
LiveBench · Reasoning
Qwen3.6 Flash leads by +23.7
gpt-oss-120b
39.2
Qwen3.6 Flash
62.9
Lmca
Qwen3.6 Flash leads by +10.4
gpt-oss-120b
26.1
Qwen3.6 Flash
36.5
Mystery Game Puzzles
Qwen3.6 Flash leads by +9.7
gpt-oss-120b
0.0
Qwen3.6 Flash
9.7
OTIS Mock AIME 2024-2025
gpt-oss-120b leads by +4.4
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
gpt-oss-120b
88.9
Qwen3.6 Flash
84.4
SimpleBench
Qwen3.6 Flash leads by +15.7
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
gpt-oss-120b
6.5
Qwen3.6 Flash
22.2
SimpleQA Verified
Qwen3.6 Flash leads by +2.0
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
gpt-oss-120b
13.9
Qwen3.6 Flash
15.9
Full benchmark table
Benchmarkgpt-oss-120bQwen3.6 Flash
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
15.815.8
Dtbench
60.561.8
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
67.777.8
LiveBench · Agentic Coding
16.746.7
LiveBench · Coding
60.264.9
LiveBench · Data Analysis
38.858.8
LiveBench · If
50.347.2
LiveBench · Language
48.663.1
LiveBench · Mathematics
68.978.9
LiveBench · Overall
46.160.4
LiveBench · Reasoning
39.262.9
Lmca
26.136.5
Mystery Game Puzzles
0.09.7
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
88.984.4
SimpleBench
SimpleBench · tests fundamental reasoning capabilities with straightforward problems designed to expose gaps in basic logical and spatial thinking.
6.522.2
SimpleQA Verified
SimpleQA Verified · short factual questions with verified answers, measuring factual accuracy and the tendency to hallucinate or provide incorrect information.
13.915.9
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
OpenAI logogpt-oss-120b$0.04$0.17131K tokens (~66 books)$0.70
Alibaba Qwen logoQwen3.6 Flash$0.19$1.131.0M tokens (~500 books)$4.22