Home/Models/Phi 2
Microsoft logo

Phi 2

by Microsoft · Released Dec 2023

Open Source
32.2
avg score
Rank #236
Compare
Better than 24% of all models
Context
N/A
Input $/1M
n/a
Output $/1M
n/a
Type
text-generation
License
Open Source
Benchmarks
14 tested
Data as of
About

Microsoft text generation model. 1759K downloads on HuggingFace.

Tested on 14 benchmarks · BenchGecko score 32.2. Top scores: ARC AI2 (67.9%), OpenBookQA (64.8%), BBH (45.9%).

Capabilities
reasoning
29.9
#119 globally
math
3.0
#268 globally
knowledge
33.9
#232 globally
general
28.0
#134 globally
language
27.4
#149 globally
Benchmark Scores
Compare All
Tested on 14 benchmarks · Ranked across 5 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
BBH

BIG-Bench Hard. 23 challenging tasks from BIG-Bench where prior language models fell below average human performance.

45.9·
MUSR

HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.

13.8·
MATH Level 5

HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.

3.0·
ARC AI2

AI2 Reasoning Challenge. Grade-school science questions requiring multi-step reasoning. Easy and Challenge sets test different difficulty levels.

67.9·
OpenBookQA

Elementary science questions with access to a small book of core science facts. Tests reasoning beyond memorization.

64.8·
TriviaQA

Trivia questions sourced from trivia enthusiasts and quiz websites. Tests breadth of general knowledge.

45.2·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Links
Documentation
Community
BenchGecko API
microsoft-phi-2
Specifications
  • Typetext-generation
  • ContextN/A
  • ReleasedDec 2023
  • LicenseOpen Source
  • StatusActive
Available On
Microsoft logoMicrosoftn/a
Share & Export
Tweet
Phi 2 is an open-source text-generation AI model by Microsoft, released in December 2023. It has an average benchmark score of 32.2.

Key facts · as of 2026-04-09

  • Phi 2 by Microsoft. BenchGecko score 32.2, rank 236 of 312 scored models (normalized average of public benchmark scores).
  • List price n/a input · n/a output per 1M tokens (as of 2026-04-09).

How to cite · data as of 2026-04-09

Phi 2 · benchmarks, pricing and providers. BenchGecko, data as of 2026-04-09. https://benchgecko.ai/model/microsoft-phi-2

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP