Compare · ModelsLive · 2 picked · head to head
PaLM 2-S vs phi-3-mini 3.8B
Side by side · benchmarks, pricing, and signals you can act on.
Winner summary
PaLM 2-S wins on 3/5 benchmarks
PaLM 2-S wins 3 of 5 shared benchmarks. Leads in knowledge.
Category leads
knowledge·PaLM 2-S
Hype vs Reality
Attention vs performance
PaLM 2-S
#53 by perf·no signal
phi-3-mini 3.8B
#86 by perf·#20 by attention
Vendor risk
Who is behind the model
U
Unknown
private · undisclosed
Microsoft
$3.84T·Big Tech
Head to head
5 benchmarks · 2 models
PaLM 2-Sphi-3-mini 3.8B
ARC AI2
phi-3-mini 3.8B leads by +33.7
AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval.
PaLM 2-S
46.1
phi-3-mini 3.8B
79.9
HellaSwag
PaLM 2-S leads by +7.1
HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.
PaLM 2-S
76.0
phi-3-mini 3.8B
68.9
OpenBookQA
phi-3-mini 3.8B leads by +42.4
OpenBookQA · science questions that require combining a given core fact with broad common knowledge, mimicking an open-book exam setting.
PaLM 2-S
41.6
phi-3-mini 3.8B
84.0
TriviaQA
PaLM 2-S leads by +11.2
TriviaQA · reading comprehension benchmark with trivia questions, requiring models to find and reason over evidence from provided documents.
PaLM 2-S
75.2
phi-3-mini 3.8B
64.0
Winogrande
PaLM 2-S leads by +14.2
WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs.
PaLM 2-S
55.8
phi-3-mini 3.8B
41.6
Full benchmark table
| Benchmark | PaLM 2-S | phi-3-mini 3.8B |
|---|---|---|
ARC AI2 AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval. | 46.1 | 79.9 |
HellaSwag HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios. | 76.0 | 68.9 |
OpenBookQA OpenBookQA · science questions that require combining a given core fact with broad common knowledge, mimicking an open-book exam setting. | 41.6 | 84.0 |
TriviaQA TriviaQA · reading comprehension benchmark with trivia questions, requiring models to find and reason over evidence from provided documents. | 75.2 | 64.0 |
Winogrande WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs. | 55.8 | 41.6 |
Pricing · per 1M tokens · projected $/mo at 10M tokens
| Model | Input | Output | Context | Projected $/mo |
|---|---|---|---|---|
U PaLM 2-S | — | — | — | — |
| — | — | — | — |