Home/Models/Phi 4 Multimodal Instruct
Microsoft logo

Phi 4 Multimodal Instruct

by Microsoft · Released Feb 2025

Open Source
Compare
Context
N/A
Input $/1M
n/a
Output $/1M
n/a
Type
automatic-speech-recognition
License
Open Source
Benchmarks
4 tested
Data as of
About

Microsoft automatic speech recognition model. 308K downloads on HuggingFace.

Tested on 4 benchmarks. Top scores: Artificial Analysis · GPQA Diamond (31.5%), Artificial Analysis · MMMU Pro (14.5%), Artificial Analysis — Quality Index (5.8%).

Capabilities
speed
14.2
#114 globally
Benchmark Scores
Compare All
Tested on 4 benchmarks · Ranked across 1 categories
Score Distribution (all 312 models)
0255075100
Artificial Analysis — Quality Index

Artificial Analysis Quality Index. Composite quality score combining multiple benchmark results into a single metric.

5.8·
Excellent (85+) Good (70-85) Average (50-70) Below (<50)
Links
Documentation
Community
BenchGecko API
microsoft-phi-4-multimodal-instruct
Specifications
  • Typeautomatic-speech-recognition
  • ContextN/A
  • ReleasedFeb 2025
  • LicenseOpen Source
  • StatusActive
Available On
Microsoft logoMicrosoftn/a
Categories
Share & Export
Tweet
Phi 4 Multimodal Instruct is an open-source automatic-speech-recognition AI model by Microsoft, released in February 2025.

Key facts · as of 2026-04-09

  • Phi 4 Multimodal Instruct by Microsoft. Not enough public benchmark scores to rank yet.
  • List price n/a input · n/a output per 1M tokens (as of 2026-04-09).

How to cite · data as of 2026-04-09

Phi 4 Multimodal Instruct · benchmarks, pricing and providers. BenchGecko, data as of 2026-04-09. https://benchgecko.ai/model/microsoft-phi-4-multimodal-instruct

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP