Tested on 12 benchmarks · BenchGecko score 43.7. Top scores: OTIS Mock AIME 2024-2025 (86.7%), GPQA diamond (79.8%), Dtbench (67.5%).
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
Original research-level math problems created by professional mathematicians. Problems are unpublished and cannot be memorized.
Hardest tier of FrontierMath. Problems at the frontier of human mathematical ability, many unsolved by most mathematicians.
Graduate-level science questions written by PhD experts. Diamond subset contains questions where experts disagree, testing deep understanding.
Simple factual questions with verified correct answers. Tests accuracy of basic knowledge retrieval. Low scores indicate hallucination.
Tactical chess puzzles testing pattern recognition and multi-move calculation. Measures strategic reasoning ability.
Agent performance evaluation testing multi-step tool use, planning, and execution in realistic environments.
- Typetext
- ContextN/A
- ReleasedJan 2024
- LicenseOpen Source
- Statusbenchmark-only
Frequently Asked Questions
Key facts · as of 2026-06-22
- Qwen 3.5 Plus (hosted 397B-A17B) by Alibaba Qwen. BenchGecko score 43.7, rank 182 of 312 scored models (normalized average of public benchmark scores).
- List price n/a input · n/a output per 1M tokens (as of 2026-06-22).
How to cite · data as of 2026-06-22
Qwen 3.5 Plus (hosted 397B-A17B) · benchmarks, pricing and providers. BenchGecko, data as of 2026-06-22. https://benchgecko.ai/model/qwen-3-5-plus-hosted-397b-a17b
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP