Home/Models/Qwen2.5 72B Instruct
Alibaba Qwen logo

Qwen2.5 72B Instruct

by Alibaba Qwen · Released Sep 2024

Open Source
59.6
avg score
Rank #90
Compare
Better than 71% of all models
Context
33K tokens (~16 books)
Input $/1M
$0.36
Output $/1M
$0.40
Type
text
License
Open Source
Benchmarks
28 tested
Data as of
About

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and...

Tested on 28 benchmarks · BenchGecko score 59.6. Top scores: Chatbot Arena Elo — Overall (1302.8%), ARC AI2 (92.7%), IFEval (86.4%).

Looking for similar performance at lower cost?
DeepSeek V4 Flash 0731 scores 60.1 (101% as good) at $0.02/1M input · 96% cheaper
Capabilities
coding
40.7
#132 globally
reasoning
42.4
#96 globally
math
43.6
#143 globally
knowledge
59.9
#57 globally
agentic
5.3
#63 globally
language
86.4
#31 globally
general
37.9
#78 globally
multimodal
64.7
#2 globally
Benchmark Scores
Compare All
Tested on 28 benchmarks · Ranked across 9 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
Aider — Code Editing

Code editing benchmark from the Aider project. Measures ability to apply targeted code changes while maintaining correctness and style.

65.4·
WeirdML

Unusual and adversarial machine learning challenges. Tests robustness of reasoning about edge cases in ML systems.

16.0·
BBH

BIG-Bench Hard. 23 challenging tasks from BIG-Bench where prior language models fell below average human performance.

73.1·
MUSR

HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.

11.7·
MATH level 5

Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.

63.2·
MATH Level 5

HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.

59.8·
OTIS Mock AIME 2024-2025

Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.

8.0·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Recently Happened
Qwen2.5 72B Instruct pricing increased 200%
Apr 27, 2026
Links
Documentation
Community
BenchGecko API
qwen-2-5-72b-instruct
Specifications
  • Typetext
  • Context33K tokens (~16 books)
  • ReleasedSep 2024
  • LicenseOpen Source
  • StatusActive
  • Cost / Message~$0.001
Available On
Alibaba Qwen logoAlibaba Qwen$0.36
Share & Export
Tweet
Qwen2.5 72B Instruct is an open-source text AI model by Alibaba Qwen, released in September 2024. It has an average benchmark score of 59.6. Context window: 33K tokens.

Key facts · as of 2026-10-05

  • Qwen2.5 72B Instruct by Alibaba Qwen. BenchGecko score 59.6, rank 89 of 312 scored models (normalized average of public benchmark scores).
  • List price $0.36 input · $0.40 output per 1M tokens (as of 2026-10-05).
  • Sold by 2 providers (as of 2026-10-05): DeepInfra (fp8) $0.36 in / $0.40 out · Novita (bf16) $0.38 in / $0.40 out. Every provider

How to cite · data as of 2026-10-05

Qwen2.5 72B Instruct · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/qwen-2-5-72b-instruct

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP