Compare · ModelsLive · 2 picked · head to head
gpt-oss-120b vs gpt-oss-20b
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
gpt-oss-120b wins on 27/29 benchmarks
gpt-oss-120b wins 27 of 29 shared benchmarks. Leads in speed · arena · knowledge.
Category leads
speed·gpt-oss-120barena·gpt-oss-120bknowledge·gpt-oss-120bgeneral·gpt-oss-120blanguage·gpt-oss-120bmath·gpt-oss-120breasoning·gpt-oss-120bcoding·gpt-oss-120b
Hype vs Reality
Attention vs performance
gpt-oss-120b
#127 by perf·no signal
gpt-oss-20b
#97 by perf·no signal
Best value
gpt-oss-20b
2.1x better value than gpt-oss-120b
gpt-oss-120b
475.4 pts/$
$0.10/M
gpt-oss-20b
983.3 pts/$
$0.05/M
Vendor risk
Who is behind the model
OpenAI
$840.0B·Tier 1
OpenAI
$840.0B·Tier 1
Head to head
29 benchmarks · 2 models
gpt-oss-120bgpt-oss-20b
Artificial Analysis · CritPt
gpt-oss-20b leads by +0.3
gpt-oss-120b
1.1
gpt-oss-20b
1.4
Artificial Analysis · GDPval
gpt-oss-120b leads by +4.8
gpt-oss-120b
4.8
gpt-oss-20b
0.0
Artificial Analysis · GPQA Diamond
gpt-oss-120b leads by +9.4
gpt-oss-120b
78.2
gpt-oss-20b
68.8
Artificial Analysis · Humanity's Last Exam
gpt-oss-120b leads by +8.6
gpt-oss-120b
19.6
gpt-oss-20b
11.0
Artificial Analysis · IFBench
gpt-oss-120b leads by +3.9
gpt-oss-120b
69.0
gpt-oss-20b
65.1
Artificial Analysis · Long Context Reasoning
gpt-oss-120b leads by +17.3
gpt-oss-120b
52.0
gpt-oss-20b
34.7
Artificial Analysis · Quality Index
gpt-oss-120b leads by +2.6
gpt-oss-120b
11.6
gpt-oss-20b
9.0
Artificial Analysis · SciCode
gpt-oss-20b leads by +4.9
gpt-oss-120b
34.0
gpt-oss-20b
38.9
Artificial Analysis · tau2-Bench Telecom
gpt-oss-120b leads by +5.6
gpt-oss-120b
65.8
gpt-oss-20b
60.2
Artificial Analysis · Terminal-Bench Hard
gpt-oss-120b leads by +12.9
gpt-oss-120b
23.5
gpt-oss-20b
10.6
Chatbot Arena Elo · Overall
gpt-oss-120b leads by +34.0
gpt-oss-120b
1351.6
gpt-oss-20b
1317.6
Chess Puzzles
gpt-oss-120b leads by +15.8
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
gpt-oss-120b
15.8
gpt-oss-20b
0.0
Dtbench
gpt-oss-120b leads by +13.8
gpt-oss-120b
60.5
gpt-oss-20b
46.7
GPQA diamond
gpt-oss-120b leads by +20.0
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
gpt-oss-120b
67.7
gpt-oss-20b
47.7
HELM · GPQA
gpt-oss-120b leads by +9.0
gpt-oss-120b
68.4
gpt-oss-20b
59.4
HELM · IFEval
gpt-oss-120b leads by +10.4
gpt-oss-120b
83.6
gpt-oss-20b
73.2
HELM · MMLU-Pro
gpt-oss-120b leads by +5.5
gpt-oss-120b
79.5
gpt-oss-20b
74.0
HELM · Omni-MATH
gpt-oss-120b leads by +12.3
gpt-oss-120b
68.8
gpt-oss-20b
56.5
HELM · WildBench
gpt-oss-120b leads by +10.8
gpt-oss-120b
84.5
gpt-oss-20b
73.7
Lmca
gpt-oss-120b leads by +9.0
gpt-oss-120b
26.1
gpt-oss-20b
17.1
OpenCompass · AIME2025
gpt-oss-120b leads by +5.5
gpt-oss-120b
93.4
gpt-oss-20b
87.9
OpenCompass · GPQA-Diamond
gpt-oss-120b leads by +10.0
gpt-oss-120b
78.9
gpt-oss-20b
68.9
OpenCompass · HLE
gpt-oss-120b leads by +6.7
gpt-oss-120b
18.3
gpt-oss-20b
11.6
OpenCompass · IFEval
gpt-oss-120b leads by +1.3
gpt-oss-120b
90.2
gpt-oss-20b
88.9
OpenCompass · LiveCodeBenchV6
gpt-oss-120b leads by +10.0
gpt-oss-120b
78.4
gpt-oss-20b
68.4
OpenCompass · MMLU-Pro
gpt-oss-120b leads by +6.9
gpt-oss-120b
79.7
gpt-oss-20b
72.8
OTIS Mock AIME 2024-2025
gpt-oss-120b leads by +23.6
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
gpt-oss-120b
88.9
gpt-oss-20b
65.2
Terminal Bench
gpt-oss-120b leads by +15.3
Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence.
gpt-oss-120b
18.7
gpt-oss-20b
3.4
WeirdML
gpt-oss-120b leads by +7.2
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
gpt-oss-120b
48.2
gpt-oss-20b
40.9
Full benchmark table
| Benchmark | gpt-oss-120b | gpt-oss-20b |
|---|---|---|
Artificial Analysis · CritPt | 1.1 | 1.4 |
Artificial Analysis · GDPval | 4.8 | 0.0 |
Artificial Analysis · GPQA Diamond | 78.2 | 68.8 |
Artificial Analysis · Humanity's Last Exam | 19.6 | 11.0 |
Artificial Analysis · IFBench | 69.0 | 65.1 |
Artificial Analysis · Long Context Reasoning | 52.0 | 34.7 |
Artificial Analysis · Quality Index | 11.6 | 9.0 |
Artificial Analysis · SciCode | 34.0 | 38.9 |
Artificial Analysis · tau2-Bench Telecom | 65.8 | 60.2 |
Artificial Analysis · Terminal-Bench Hard | 23.5 | 10.6 |
Chatbot Arena Elo · Overall | 1351.6 | 1317.6 |
Chess Puzzles Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities. | 15.8 | 0.0 |
Dtbench | 60.5 | 46.7 |
GPQA diamond Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs. | 67.7 | 47.7 |
HELM · GPQA | 68.4 | 59.4 |
HELM · IFEval | 83.6 | 73.2 |
HELM · MMLU-Pro | 79.5 | 74.0 |
HELM · Omni-MATH | 68.8 | 56.5 |
HELM · WildBench | 84.5 | 73.7 |
Lmca | 26.1 | 17.1 |
OpenCompass · AIME2025 | 93.4 | 87.9 |
OpenCompass · GPQA-Diamond | 78.9 | 68.9 |
OpenCompass · HLE | 18.3 | 11.6 |
OpenCompass · IFEval | 90.2 | 88.9 |
OpenCompass · LiveCodeBenchV6 | 78.4 | 68.4 |
OpenCompass · MMLU-Pro | 79.7 | 72.8 |
OTIS Mock AIME 2024-2025 OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills. | 88.9 | 65.2 |
Terminal Bench Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence. | 18.7 | 3.4 |
WeirdML WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns. | 48.2 | 40.9 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.04 | $0.17 | 131K tokens (~66 books) | $0.70 | |
| $0.02 | $0.09 | 131K tokens (~66 books) | $0.36 |