Microsoft Research Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed. At 14 billion...
Tested on 24 benchmarks · BenchGecko score 49.0. Top scores: Chatbot Arena Elo — Overall (1256.3%), MMLU (79.7%), MATH level 5 (64.9%).
HuggingFace MuSR (Multi-Step Reasoning). Tests multi-hop reasoning requiring chaining multiple facts together.
Competition-level math from AMC, AIME, and olympiad problems. Level 5 is the hardest tier, requiring creative problem-solving.
HuggingFace evaluation of MATH Level 5 problems. Competition math requiring advanced reasoning and proof construction.
Mock AIME (American Invitational Mathematics Exam) problems from OTIS. Tests mathematical competition performance.
Massive Multitask Language Understanding. 57 subjects from STEM, humanities, and social sciences. The most widely-cited knowledge benchmark.
Writing quality evaluation by Lech Mazur. Tests prose quality, coherence, and stylistic ability.
HuggingFace MMLU-Pro. Harder version of MMLU with 10 answer choices instead of 4 and more challenging questions.
- Typetext
- Context16K tokens (~8 books)
- ReleasedJan 2025
- LicenseOpen Source
- StatusActive
- Cost / Message~$0.000
Frequently Asked Questions
Key facts · as of 2026-10-05
- Phi 4 by Microsoft. BenchGecko score 49.0, rank 151 of 312 scored models (normalized average of public benchmark scores).
- List price $0.0700 input · $0.14 output per 1M tokens (as of 2026-10-05).
- Sold by 1 provider (as of 2026-10-05): DeepInfra (bf16) $0.0700 in / $0.14 out.
How to cite · data as of 2026-10-05
Phi 4 · benchmarks, pricing and providers. BenchGecko, data as of 2026-10-05. https://benchgecko.ai/model/phi-4
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP