Deepseek-ai text generation model. 958K downloads on HuggingFace.
Tested on 15 benchmarks · BenchGecko score 58.8. Top scores: JCommonsenseQA (95.3%), JSQuAD (89.2%), JNLI (85.6%).
HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.
Graduate-level science questions written by PhD experts. Diamond subset contains questions where experts disagree, testing deep understanding.
HuggingFace MMLU-Pro. Harder version of MMLU with 10 answer choices instead of 4 and more challenging questions.
Broad Assessment of Language and Reasoning Over Games. Tests strategic and logical reasoning through game scenarios.
- Typetext-generation
- ContextN/A
- ReleasedJan 2025
- LicenseOpen Source
- StatusActive
Frequently Asked Questions
Key facts · as of 2026-06-22
- DeepSeek R1 Distill Qwen 32B by DeepSeek. BenchGecko score 58.8, rank 93 of 312 scored models (normalized average of public benchmark scores).
- List price n/a input · n/a output per 1M tokens (as of 2026-06-22).
How to cite · data as of 2026-06-22
DeepSeek R1 Distill Qwen 32B · benchmarks, pricing and providers. BenchGecko, data as of 2026-06-22. https://benchgecko.ai/model/deepseek-ai-deepseek-r1-distill-qwen-32b
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP