Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in **code generation**, **code reasoning**...
Tested on 14 benchmarks · BenchGecko score 71.4. Top scores: Chatbot Arena Elo — Overall (1270.6%), GSM8K (91.1%), HellaSwag (77.3%).
MiniMax M2 scores 72.4 (101% as good) at $0.30/1M input · 55% cheaper
Code editing benchmark from the Aider project. Measures ability to apply targeted code changes while maintaining correctness and style.
Multi-language code editing from Aider. Tests editing ability across Python, JavaScript, TypeScript, Java, C++, Go, Rust, and more.
HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.
Grade school math word problems. 8,500 problems testing multi-step arithmetic reasoning. A foundational math benchmark.
HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.
- Typetext
- Context33K tokens (~16 books)
- ReleasedNov 2024
- LicenseOpen Source
- StatusActive
- Cost / Message~$0.002
Frequently Asked Questions
Key facts · as of 2026-10-05
- Qwen2.5 Coder 32B Instruct by Alibaba Qwen. BenchGecko score 71.4, rank 34 of 312 scored models (normalized average of public benchmark scores).
- List price $0.66 input · $1.00 output per 1M tokens (as of 2026-10-05).
- Sold by 1 provider (as of 2026-10-05): Cloudflare $0.66 in / $1.00 out.
How to cite · data as of 2026-10-05
Qwen2.5 Coder 32B Instruct · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/qwen-2-5-coder-32b-instruct
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP