Compare · ModelsLive · 2 picked · head to head

Qwen2.5 72B Instruct vs Stable Beluga 2

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

Qwen2.5 72B Instruct wins 9 of 11 shared benchmarks. Leads in knowledge · reasoning · general.

Category leads
knowledge·Qwen2.5 72B Instructreasoning·Qwen2.5 72B Instructgeneral·Qwen2.5 72B Instructlanguage·Qwen2.5 72B Instructmath·Qwen2.5 72B Instruct
Hype vs Reality
Qwen2.5 72B Instruct
#128 by perf·#2 by attention
DESERVED
Stable Beluga 2
#141 by perf·no signal
QUIET
Best value
Qwen2.5 72B Instruct
129.5 pts/$
$0.38/M
Stable Beluga 2
n/a
no price
Vendor risk
Alibaba Qwen logo
Alibaba (Qwen)
$293.0B·Tier 1
Low risk
Unknown
private · undisclosed
Unknown
Head to head
Qwen2.5 72B InstructStable Beluga 2
ARC AI2
Qwen2.5 72B Instruct leads by +11.2
AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval.
Qwen2.5 72B Instruct
92.7
Stable Beluga 2
81.5
BBH
Qwen2.5 72B Instruct leads by +14.0
BIG-Bench Hard · a curated subset of 23 challenging tasks from BIG-Bench where language models previously failed to outperform average humans.
Qwen2.5 72B Instruct
73.1
Stable Beluga 2
59.1
HellaSwag
Qwen2.5 72B Instruct leads by +0.9
HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.
Qwen2.5 72B Instruct
79.7
Stable Beluga 2
78.8
BBH (HuggingFace)
Qwen2.5 72B Instruct leads by +20.6
Qwen2.5 72B Instruct
61.9
Stable Beluga 2
41.3
GPQA
Qwen2.5 72B Instruct leads by +7.8
Qwen2.5 72B Instruct
16.7
Stable Beluga 2
8.8
IFEval
Qwen2.5 72B Instruct leads by +48.5
Qwen2.5 72B Instruct
86.4
Stable Beluga 2
37.9
MATH Level 5
Qwen2.5 72B Instruct leads by +55.4
Qwen2.5 72B Instruct
59.8
Stable Beluga 2
4.4
MMLU-PRO
Qwen2.5 72B Instruct leads by +25.5
Qwen2.5 72B Instruct
51.4
Stable Beluga 2
25.9
MUSR
Stable Beluga 2 leads by +6.9
Qwen2.5 72B Instruct
11.7
Stable Beluga 2
18.6
MMLU
Qwen2.5 72B Instruct leads by +22.3
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Qwen2.5 72B Instruct
80.4
Stable Beluga 2
58.1
PIQA
Stable Beluga 2 leads by +1.4
PIQA (Physical Interaction QA) · tests intuitive physical reasoning by asking models to select the correct approach for everyday physical tasks.
Qwen2.5 72B Instruct
65.2
Stable Beluga 2
66.6
Full benchmark table
BenchmarkQwen2.5 72B InstructStable Beluga 2
ARC AI2
AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval.
92.781.5
BBH
BIG-Bench Hard · a curated subset of 23 challenging tasks from BIG-Bench where language models previously failed to outperform average humans.
73.159.1
HellaSwag
HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.
79.778.8
BBH (HuggingFace)
61.941.3
GPQA
16.78.8
IFEval
86.437.9
MATH Level 5
59.84.4
MMLU-PRO
51.425.9
MUSR
11.718.6
MMLU
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
80.458.1
PIQA
PIQA (Physical Interaction QA) · tests intuitive physical reasoning by asking models to select the correct approach for everyday physical tasks.
65.266.6
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
Alibaba Qwen logoQwen2.5 72B Instruct$0.36$0.4033K tokens (~16 books)$3.70
Stable Beluga 2————