Tested on 11 benchmarks · BenchGecko score 17.2. Top scores: Winogrande (46.8%), HellaSwag (30.1%), ARC AI2 (25.9%).
HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.
HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.
Commonsense coreference resolution. Tests understanding of pronoun references in ambiguous sentences.
Sentence completion requiring commonsense reasoning about physical and social situations. Tests real-world understanding.
AI2 Reasoning Challenge. Grade-school science questions requiring multi-step reasoning. Easy and Challenge sets test different difficulty levels.
- Typetext
- ContextN/A
- ReleasedJan 2024
- LicenseOpen Source
- Statusbenchmark-only
Frequently Asked Questions
Key facts · as of 2026-04-09
- Phi-1.5 by Microsoft. BenchGecko score 17.2, rank 287 of 312 scored models (normalized average of public benchmark scores).
- List price n/a input · n/a output per 1M tokens (as of 2026-04-09).
How to cite · data as of 2026-04-09
Phi-1.5 · benchmarks, pricing and providers. BenchGecko, data as of 2026-04-09. https://benchgecko.ai/model/phi-1-5
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP