Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
Tested on 12 benchmarks · BenchGecko score 36.3. Top scores: OpenCompass — IFEval (85.6%), OpenCompass — MMLU-Pro (72.1%), OpenCompass — AIME2025 (66.2%).
Qwen3.5-Flash scores 36.5 (101% as good) at $0.07/1M input · 44% cheaper
OpenCompass Live Code Bench v6. Fresh competitive programming problems to evaluate code generation without memorization.
OpenCompass evaluation on AIME 2025 problems. Tests mathematical reasoning on fresh competition problems.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
OpenCompass MMLU-Pro evaluation. Harder knowledge test with more answer choices.
LiveBench fiction analysis. Tests literary comprehension and creative text understanding.
OpenCompass evaluation of GPQA Diamond. PhD-level science questions from the hardest subset.
- Typetext
- Context131K tokens (~66 books)
- ReleasedApr 2025
- LicenseOpen Source
- StatusActive
- Cost / Message~$0.001
Frequently Asked Questions
Related Models
Llama 3.2 90B · MetaGemini 1.5 Pro (May 2024) · Google DeepMindQwen2.5-Max · Alibaba QwenLlama 2-13B · MetaQwen3.5-Flash · Alibaba QwenKey facts · as of 2026-10-05
- Qwen3 8B by Alibaba Qwen. BenchGecko score 36.3, rank 215 of 312 scored models (normalized average of public benchmark scores).
- List price $0.12 input · $0.46 output per 1M tokens (as of 2026-10-05).
- Sold by 1 provider (as of 2026-10-05): Alibaba $0.12 in / $0.46 out.
How to cite · data as of 2026-10-05
Qwen3 8B · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/qwen3-8b
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP