Compare · ModelsLive · 2 picked · head to head

gpt-oss-120b vs gpt-oss-20b

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

gpt-oss-120b wins 27 of 29 shared benchmarks. Leads in speed · arena · knowledge.

Category leads
speed·gpt-oss-120barena·gpt-oss-120bknowledge·gpt-oss-120bgeneral·gpt-oss-120blanguage·gpt-oss-120bmath·gpt-oss-120breasoning·gpt-oss-120bcoding·gpt-oss-120b
Hype vs Reality
gpt-oss-120b
#127 by perf·no signal
QUIET
gpt-oss-20b
#97 by perf·no signal
QUIET
Best value
2.1x better value than gpt-oss-120b
gpt-oss-120b
475.4 pts/$
$0.10/M
gpt-oss-20b
983.3 pts/$
$0.05/M
Vendor risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
OpenAI logo
OpenAI
$840.0B·Tier 1
Medium risk
Head to head
gpt-oss-120bgpt-oss-20b
Artificial Analysis · CritPt
gpt-oss-20b leads by +0.3
gpt-oss-120b
1.1
gpt-oss-20b
1.4
Artificial Analysis · GDPval
gpt-oss-120b leads by +4.8
gpt-oss-120b
4.8
gpt-oss-20b
0.0
Artificial Analysis · GPQA Diamond
gpt-oss-120b leads by +9.4
gpt-oss-120b
78.2
gpt-oss-20b
68.8
Artificial Analysis · Humanity's Last Exam
gpt-oss-120b leads by +8.6
gpt-oss-120b
19.6
gpt-oss-20b
11.0
Artificial Analysis · IFBench
gpt-oss-120b leads by +3.9
gpt-oss-120b
69.0
gpt-oss-20b
65.1
Artificial Analysis · Long Context Reasoning
gpt-oss-120b leads by +17.3
gpt-oss-120b
52.0
gpt-oss-20b
34.7
Artificial Analysis · Quality Index
gpt-oss-120b leads by +2.6
gpt-oss-120b
11.6
gpt-oss-20b
9.0
Artificial Analysis · SciCode
gpt-oss-20b leads by +4.9
gpt-oss-120b
34.0
gpt-oss-20b
38.9
Artificial Analysis · tau2-Bench Telecom
gpt-oss-120b leads by +5.6
gpt-oss-120b
65.8
gpt-oss-20b
60.2
Artificial Analysis · Terminal-Bench Hard
gpt-oss-120b leads by +12.9
gpt-oss-120b
23.5
gpt-oss-20b
10.6
Chatbot Arena Elo · Overall
gpt-oss-120b leads by +34.0
gpt-oss-120b
1351.6
gpt-oss-20b
1317.6
Chess Puzzles
gpt-oss-120b leads by +15.8
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
gpt-oss-120b
15.8
gpt-oss-20b
0.0
Dtbench
gpt-oss-120b leads by +13.8
gpt-oss-120b
60.5
gpt-oss-20b
46.7
GPQA diamond
gpt-oss-120b leads by +20.0
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
gpt-oss-120b
67.7
gpt-oss-20b
47.7
HELM · GPQA
gpt-oss-120b leads by +9.0
gpt-oss-120b
68.4
gpt-oss-20b
59.4
HELM · IFEval
gpt-oss-120b leads by +10.4
gpt-oss-120b
83.6
gpt-oss-20b
73.2
HELM · MMLU-Pro
gpt-oss-120b leads by +5.5
gpt-oss-120b
79.5
gpt-oss-20b
74.0
HELM · Omni-MATH
gpt-oss-120b leads by +12.3
gpt-oss-120b
68.8
gpt-oss-20b
56.5
HELM · WildBench
gpt-oss-120b leads by +10.8
gpt-oss-120b
84.5
gpt-oss-20b
73.7
Lmca
gpt-oss-120b leads by +9.0
gpt-oss-120b
26.1
gpt-oss-20b
17.1
OpenCompass · AIME2025
gpt-oss-120b leads by +5.5
gpt-oss-120b
93.4
gpt-oss-20b
87.9
OpenCompass · GPQA-Diamond
gpt-oss-120b leads by +10.0
gpt-oss-120b
78.9
gpt-oss-20b
68.9
OpenCompass · HLE
gpt-oss-120b leads by +6.7
gpt-oss-120b
18.3
gpt-oss-20b
11.6
OpenCompass · IFEval
gpt-oss-120b leads by +1.3
gpt-oss-120b
90.2
gpt-oss-20b
88.9
OpenCompass · LiveCodeBenchV6
gpt-oss-120b leads by +10.0
gpt-oss-120b
78.4
gpt-oss-20b
68.4
OpenCompass · MMLU-Pro
gpt-oss-120b leads by +6.9
gpt-oss-120b
79.7
gpt-oss-20b
72.8
OTIS Mock AIME 2024-2025
gpt-oss-120b leads by +23.6
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
gpt-oss-120b
88.9
gpt-oss-20b
65.2
Terminal Bench
gpt-oss-120b leads by +15.3
Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence.
gpt-oss-120b
18.7
gpt-oss-20b
3.4
WeirdML
gpt-oss-120b leads by +7.2
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
gpt-oss-120b
48.2
gpt-oss-20b
40.9
Full benchmark table
Benchmarkgpt-oss-120bgpt-oss-20b
Artificial Analysis · CritPt
1.11.4
Artificial Analysis · GDPval
4.80.0
Artificial Analysis · GPQA Diamond
78.268.8
Artificial Analysis · Humanity's Last Exam
19.611.0
Artificial Analysis · IFBench
69.065.1
Artificial Analysis · Long Context Reasoning
52.034.7
Artificial Analysis · Quality Index
11.69.0
Artificial Analysis · SciCode
34.038.9
Artificial Analysis · tau2-Bench Telecom
65.860.2
Artificial Analysis · Terminal-Bench Hard
23.510.6
Chatbot Arena Elo · Overall
1351.61317.6
Chess Puzzles
Chess Puzzles · tests strategic and tactical reasoning by having models solve chess puzzle positions, evaluating lookahead and pattern recognition abilities.
15.80.0
Dtbench
60.546.7
GPQA diamond
Graduate-Level Google-Proof QA (Diamond set) · expert-crafted questions in physics, biology, and chemistry that are difficult even for domain PhDs.
67.747.7
HELM · GPQA
68.459.4
HELM · IFEval
83.673.2
HELM · MMLU-Pro
79.574.0
HELM · Omni-MATH
68.856.5
HELM · WildBench
84.573.7
Lmca
26.117.1
OpenCompass · AIME2025
93.487.9
OpenCompass · GPQA-Diamond
78.968.9
OpenCompass · HLE
18.311.6
OpenCompass · IFEval
90.288.9
OpenCompass · LiveCodeBenchV6
78.468.4
OpenCompass · MMLU-Pro
79.772.8
OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 · simulated American Invitational Mathematics Examination problems testing advanced problem-solving skills.
88.965.2
Terminal Bench
Terminal-Bench 2.0 · evaluates AI agents on real terminal-based coding tasks · writing scripts, debugging, running tests, and managing projects entirely through command-line interaction. Tests both code quality and terminal fluency. Claude Opus 4.7 scores 69.4%, demonstrating significant agentic terminal competence.
18.73.4
WeirdML
WeirdML · tests models on unusual and adversarial machine learning tasks that require creative problem-solving beyond standard patterns.
48.240.9
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
OpenAI logogpt-oss-120b$0.04$0.17131K tokens (~66 books)$0.70
OpenAI logogpt-oss-20b$0.02$0.09131K tokens (~66 books)$0.36