Home/Models/Phi 4
Microsoft logo

Phi 4

by Microsoft · Released Jan 2025

Open Source
49.0
avg score
Rank #151
Compare
Better than 52% of all models
Context
16K tokens (~8 books)
Input $/1M
$0.07
Output $/1M
$0.14
Type
text
License
Open Source
Benchmarks
24 tested
Data as of
About

Microsoft Research Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...

Tested on 24 benchmarks · BenchGecko score 49.0. Top scores: Chatbot Arena Elo — Overall (1256.3%), MMLU (79.7%), MATH level 5 (64.9%).

Capabilities
reasoning
11.4
#166 globally
math
41.3
#150 globally
knowledge
36.1
#219 globally
speed
10.6
#122 globally
language
64.8
#95 globally
general
55.7
#21 globally
Benchmark Scores
Compare All
Tested on 24 benchmarks · Ranked across 7 categories
Score Distribution (all 312 models)
0255075100
▲ You are here
MUSR

HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.

11.4·
MATH level 5

Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.

64.9·
MATH Level 5

HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.

45.2·
OTIS Mock AIME 2024-2025

Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.

13.7·
MMLU

Massive Multitask Language Understanding. 57 subjects from STEM, humanities, and social sciences. The most widely-cited knowledge benchmark.

79.7·
Lech Mazur Writing

Writing quality evaluation by Lech Mazur. Tests prose quality, coherence, and stylistic ability.

62.6·
MMLU-PRO

HuggingFace MMLU-Pro. Harder version of MMLU with 10 answer choices instead of 4 and more challenging questions.

48.3·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Recently Happened
Phi 4 pricing increased 8%
Jun 24, 2026
Specifications
  • Typetext
  • Context16K tokens (~8 books)
  • ReleasedJan 2025
  • LicenseOpen Source
  • StatusActive
  • Cost / Message~$0.000
Available On
Microsoft logoMicrosoft$0.07
Share & Export
Tweet
Phi 4 is an open-source text AI model by Microsoft, released in January 2025. It has an average benchmark score of 49.0. Context window: 16K tokens.

Key facts · as of 2026-10-05

  • Phi 4 by Microsoft. BenchGecko score 49.0, rank 151 of 312 scored models (normalized average of public benchmark scores).
  • List price $0.0700 input · $0.14 output per 1M tokens (as of 2026-10-05).
  • Sold by 1 provider (as of 2026-10-05): DeepInfra (bf16) $0.0700 in / $0.14 out.

How to cite · data as of 2026-10-05

Phi 4 · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/phi-4

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP