Tested on 11 benchmarks · BenchGecko score 73.3. Top scores: Chatbot Arena Elo — Overall (1489.3%), OTIS Mock AIME 2024-2025 (88.9%), GPQA diamond (86.4%).
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
Original research-level math problems created by professional mathematicians. Problems are unpublished and cannot be memorized.
Hardest tier of FrontierMath. Problems at the frontier of human mathematical ability, many unsolved by most mathematicians.
Graduate-level science questions written by PhD experts. Diamond subset contains questions where experts disagree, testing deep understanding.
Simple factual questions with verified correct answers. Tests accuracy of basic knowledge retrieval. Low scores indicate hallucination.
Humanitys Last Exam. Expert-level questions spanning all academic disciplines, designed to be the hardest knowledge test for AI.
Chatbot Arena overall Elo rating. Crowdsourced human preference ranking from blind head-to-head comparisons across all topics.
- Typetext
- ContextN/A
- ReleasedJan 2024
- LicenseProprietary
- Statusbenchmark-only
Frequently Asked Questions
Related Models
Hermes 3 70B Instruct · nousresearchGemini 2.5 Pro Preview 06-05 · Google DeepMindDeepSeek R1 Distill Qwen 14B · DeepSeekMiniMax M2 · minimaxgpt-oss-120b (free) · OpenAIKey facts · as of 2026-04-09
- Muse Spark by Unknown. BenchGecko score 73.3, rank 29 of 312 scored models (normalized average of public benchmark scores).
- List price n/a input · n/a output per 1M tokens (as of 2026-04-09).
How to cite · data as of 2026-04-09
Muse Spark · benchmarks, pricing and providers. BenchGecko, data as of 2026-04-09. https://benchgecko.ai/model/muse-spark
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP