Compare · ModelsLive · 2 picked · head to head
Qwen2.5 72B Instruct vs Stable Beluga 2
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
Qwen2.5 72B Instruct wins on 9/11 benchmarks
Qwen2.5 72B Instruct wins 9 of 11 shared benchmarks. Leads in knowledge · reasoning · general.
Category leads
knowledge·Qwen2.5 72B Instructreasoning·Qwen2.5 72B Instructgeneral·Qwen2.5 72B Instructlanguage·Qwen2.5 72B Instructmath·Qwen2.5 72B Instruct
Hype vs Reality
Attention vs performance
Qwen2.5 72B Instruct
#128 by perf·#2 by attention
Stable Beluga 2
#141 by perf·no signal
Best value
Qwen2.5 72B Instruct
Qwen2.5 72B Instruct
129.5 pts/$
$0.38/M
Stable Beluga 2
n/a
no price
Vendor risk
Who is behind the model
Alibaba (Qwen)
$293.0B·Tier 1
U
Unknown
private · undisclosed
Head to head
11 benchmarks · 2 models
Qwen2.5 72B InstructStable Beluga 2
ARC AI2
Qwen2.5 72B Instruct leads by +11.2
AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval.
Qwen2.5 72B Instruct
92.7
Stable Beluga 2
81.5
BBH
Qwen2.5 72B Instruct leads by +14.0
BIG-Bench Hard · a curated subset of 23 challenging tasks from BIG-Bench where language models previously failed to outperform average humans.
Qwen2.5 72B Instruct
73.1
Stable Beluga 2
59.1
HellaSwag
Qwen2.5 72B Instruct leads by +0.9
HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.
Qwen2.5 72B Instruct
79.7
Stable Beluga 2
78.8
BBH (HuggingFace)
Qwen2.5 72B Instruct leads by +20.6
Qwen2.5 72B Instruct
61.9
Stable Beluga 2
41.3
GPQA
Qwen2.5 72B Instruct leads by +7.8
Qwen2.5 72B Instruct
16.7
Stable Beluga 2
8.8
IFEval
Qwen2.5 72B Instruct leads by +48.5
Qwen2.5 72B Instruct
86.4
Stable Beluga 2
37.9
MATH Level 5
Qwen2.5 72B Instruct leads by +55.4
Qwen2.5 72B Instruct
59.8
Stable Beluga 2
4.4
MMLU-PRO
Qwen2.5 72B Instruct leads by +25.5
Qwen2.5 72B Instruct
51.4
Stable Beluga 2
25.9
MUSR
Stable Beluga 2 leads by +6.9
Qwen2.5 72B Instruct
11.7
Stable Beluga 2
18.6
MMLU
Qwen2.5 72B Instruct leads by +22.3
Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge.
Qwen2.5 72B Instruct
80.4
Stable Beluga 2
58.1
PIQA
Stable Beluga 2 leads by +1.4
PIQA (Physical Interaction QA) · tests intuitive physical reasoning by asking models to select the correct approach for everyday physical tasks.
Qwen2.5 72B Instruct
65.2
Stable Beluga 2
66.6
Full benchmark table
| Benchmark | Qwen2.5 72B Instruct | Stable Beluga 2 |
|---|---|---|
ARC AI2 AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval. | 92.7 | 81.5 |
BBH BIG-Bench Hard · a curated subset of 23 challenging tasks from BIG-Bench where language models previously failed to outperform average humans. | 73.1 | 59.1 |
HellaSwag HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios. | 79.7 | 78.8 |
BBH (HuggingFace) | 61.9 | 41.3 |
GPQA | 16.7 | 8.8 |
IFEval | 86.4 | 37.9 |
MATH Level 5 | 59.8 | 4.4 |
MMLU-PRO | 51.4 | 25.9 |
MUSR | 11.7 | 18.6 |
MMLU Massive Multitask Language Understanding · 57 subjects spanning STEM, humanities, social sciences, and more. The standard benchmark for broad knowledge. | 80.4 | 58.1 |
PIQA PIQA (Physical Interaction QA) · tests intuitive physical reasoning by asking models to select the correct approach for everyday physical tasks. | 65.2 | 66.6 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
| $0.36 | $0.40 | 33K tokens (~16 books) | $3.70 | |
U Stable Beluga 2 | — | — | — | — |