Compare · ModelsLive · 2 picked · head to head

PaLM 2-S vs phi-3-mini 3.8B

Side by side · benchmarks, pricing, and signals you can act on.

Winner summary

PaLM 2-S wins 3 of 5 shared benchmarks. Leads in knowledge.

Category leads
knowledge·PaLM 2-S
Hype vs Reality
PaLM 2-S
#53 by perf·no signal
QUIET
phi-3-mini 3.8B
#86 by perf·#20 by attention
UNDERRATED
Best value
PaLM 2-S
n/a
no price
phi-3-mini 3.8B
n/a
no price
Vendor risk
Unknown
private · undisclosed
Unknown
Microsoft logo
Microsoft
$3.84T·Big Tech
Low risk
Head to head
PaLM 2-Sphi-3-mini 3.8B
ARC AI2
phi-3-mini 3.8B leads by +33.7
AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval.
PaLM 2-S
46.1
phi-3-mini 3.8B
79.9
HellaSwag
PaLM 2-S leads by +7.1
HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.
PaLM 2-S
76.0
phi-3-mini 3.8B
68.9
OpenBookQA
phi-3-mini 3.8B leads by +42.4
OpenBookQA · science questions that require combining a given core fact with broad common knowledge, mimicking an open-book exam setting.
PaLM 2-S
41.6
phi-3-mini 3.8B
84.0
TriviaQA
PaLM 2-S leads by +11.2
TriviaQA · reading comprehension benchmark with trivia questions, requiring models to find and reason over evidence from provided documents.
PaLM 2-S
75.2
phi-3-mini 3.8B
64.0
Winogrande
PaLM 2-S leads by +14.2
WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs.
PaLM 2-S
55.8
phi-3-mini 3.8B
41.6
Full benchmark table
BenchmarkPaLM 2-Sphi-3-mini 3.8B
ARC AI2
AI2 Reasoning Challenge · tests grade-school level science knowledge with multiple-choice questions requiring reasoning beyond simple retrieval.
46.179.9
HellaSwag
HellaSwag · tests commonsense reasoning by asking models to predict the most plausible continuation of everyday scenarios.
76.068.9
OpenBookQA
OpenBookQA · science questions that require combining a given core fact with broad common knowledge, mimicking an open-book exam setting.
41.684.0
TriviaQA
TriviaQA · reading comprehension benchmark with trivia questions, requiring models to find and reason over evidence from provided documents.
75.264.0
Winogrande
WinoGrande · large-scale commonsense reasoning benchmark where models must resolve ambiguous pronouns in carefully constructed sentence pairs.
55.841.6
Pricing · per 1M tokens · projected $/mo at 10M tokens
ModelInputOutputContextProjected $/mo
PaLM 2-S————
Microsoft logophi-3-mini 3.8B————