Home/Models/Qwen2.5 32B Instruct
Alibaba logo

Qwen2.5 32B Instruct

by Alibaba · Released Sep 2024

Open Source
57.9
avg score
Rank #101
Compare
Better than 68% of all models
Context
N/A
Input $/1M
n/a
Output $/1M
n/a
Type
text-generation
License
Open Source
Benchmarks
11 tested
Data as of
About

Qwen text generation model. 3882K downloads on HuggingFace.

Tested on 11 benchmarks · BenchGecko score 57.9. Top scores: IFEval (83.5%), MATH Level 5 (62.5%), BBH (HuggingFace) (56.5%).

Capabilities
reasoning
13.5
#156 globally
math
42.0
#147 globally
knowledge
22.9
#256 globally
language
83.5
#45 globally
general
56.5
#20 globally
safety
22.9
#4 globally
Benchmark Scores
Compare All
Tested on 11 benchmarks · Ranked across 6 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
MUSR

HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.

13.5·
MATH Level 5

HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.

62.5·
MATH level 5

Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.

56.1·
OTIS Mock AIME 2024-2025

Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.

7.3·
MMLU-PRO

HuggingFace MMLU-Pro. Harder version of MMLU with 10 answer choices instead of 4 and more challenging questions.

51.9·
GPQA diamond

Graduate-level science questions written by PhD experts. Diamond subset contains questions where experts disagree, testing deep understanding.

28.1·
GPQA

HuggingFace evaluation of GPQA (Graduate-Level Google-Proof Q&A). PhD-level science questions that cannot be easily searched.

11.7·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Links
Documentation
Community
BenchGecko API
qwen-qwen25-32b-instruct
Specifications
  • Typetext-generation
  • ContextN/A
  • ReleasedSep 2024
  • LicenseOpen Source
  • StatusActive
Available On
Alibaba logoAlibaban/a
Share & Export
Tweet
Qwen2.5 32B Instruct is an open-source text-generation AI model by Alibaba, released in September 2024. It has an average benchmark score of 57.9.

Key facts · as of 2026-04-09

  • Qwen2.5 32B Instruct by Alibaba. BenchGecko score 57.9, rank 101 of 312 scored models (normalized average of public benchmark scores).
  • List price n/a input · n/a output per 1M tokens (as of 2026-04-09).

How to cite · data as of 2026-04-09

Qwen2.5 32B Instruct · benchmarks, pricing and providers. BenchGecko, data as of 2026-04-09. https://benchgecko.ai/model/qwen-qwen25-32b-instruct

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP