Home/Models/gpt-oss-120b
OpenAI logo

gpt-oss-120b

by OpenAI · Released Aug 2025

Open Source
46.7
avg score
Rank #164
Compare
Better than 47% of all models
Context
131K tokens (~66 books)
Input $/1M
$0.04
Output $/1M
$0.17
Type
text
License
Open Source
Benchmarks
48 tested
Data as of
About

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

Tested on 48 benchmarks · BenchGecko score 46.7. Top scores: Chatbot Arena Elo — Overall (1351.6%), OpenCompass — AIME2025 (93.4%), OpenCompass — IFEval (90.2%).

Capabilities
coding
41.4
#130 globally
reasoning
42.3
#97 globally
math
80.0
#26 globally
knowledge
53.6
#98 globally
agentic
4.7
#65 globally
general
33.6
#112 globally
speed
36.0
#73 globally
safety
8.2
#7 globally
language
68.2
#84 globally
Benchmark Scores
Compare All
Tested on 48 benchmarks · Ranked across 10 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
OpenCompass — LiveCodeBenchV6

OpenCompass Live Code Bench v6. Fresh competitive programming problems to evaluate code generation without memorization.

78.4·
LiveBench — Coding

Regularly refreshed coding problems that avoid data contamination. New problems added monthly to prevent memorization.

60.2·
WeirdML

Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.

48.2·
HELM — WildBench

Stanford HELM WildBench evaluation. Tests reasoning on challenging real-world tasks.

84.5·
LiveBench — Reasoning

Regularly refreshed reasoning problems testing logical deduction, spatial reasoning, and analytical thinking.

39.2·
LiveBench — Data Analysis

Fresh data analysis tasks testing ability to interpret tables, charts, and statistical data.

38.8·
OpenCompass — AIME2025

OpenCompass evaluation on AIME 2025 problems. Tests mathematical reasoning on fresh competition problems.

93.4·
OTIS Mock AIME 2024-2025

Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.

88.9·
LiveBench — Mathematics

Regularly updated math problems that test numerical reasoning, algebra, calculus, and combinatorics.

68.9·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Recently Happened
gpt-oss-120b pricing dropped 75%
Sep 30, 2026
gpt-oss-120b pricing increased 305%
Sep 18, 2026
gpt-oss-120b pricing increased 23%
Aug 23, 2026
gpt-oss-120b pricing increased 20%
Jul 10, 2026
gpt-oss-120b pricing dropped 23%
Jun 26, 2026
gpt-oss-120b pricing dropped 5%
May 1, 2026
Links
Documentation
Community
BenchGecko API
gpt-oss-120b
Specifications
  • Typetext
  • Context131K tokens (~66 books)
  • ReleasedAug 2025
  • LicenseOpen Source
  • StatusActive
  • Cost / Message~$0.000
Available On
OpenAI logoOpenAI$0.04
Share & Export
Tweet
gpt-oss-120b is an open-source text AI model by OpenAI, released in August 2025. It has an average benchmark score of 46.7. Context window: 131K tokens.

Key facts · as of 2026-10-05

  • gpt-oss-120b by OpenAI. BenchGecko score 46.7, rank 164 of 312 scored models (normalized average of public benchmark scores).
  • List price $0.0370 input · $0.17 output per 1M tokens (as of 2026-10-05).
  • Sold by 22 providers (as of 2026-10-05): CoreWeave (fp4) $0.0300 in / $0.17 out · DekaLLM (bf16) $0.0300 in / $0.18 out · AkashML (bf16) $0.0370 in / $0.19 out · DeepInfra (bf16) $0.0370 in / $0.17 out · Mancer 2 (fp8) $0.0450 in / $0.25 out · and 17 more. Every provider

How to cite · data as of 2026-10-05

gpt-oss-120b · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/gpt-oss-120b

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP