MiMo-V2-Omni is a frontier omni-modal model that natively processes image, video, and audio inputs within a unified architecture. It combines strong multimodal perception with agentic capability - visual grounding, multi-step...
Tested on 3 benchmarks. Top scores: Artificial Analysis — Agentic Index (58.6%), Artificial Analysis — Quality Index (43.4%), Artificial Analysis — Coding Index (35.5%).
Artificial Analysis Agentic Index. Composite score measuring agent capability across tool use and planning tasks.
Artificial Analysis Quality Index. Composite quality score combining multiple benchmark results into a single metric.
Artificial Analysis Coding Index. Composite coding quality score from multiple code benchmarks.
- Typemultimodal
- Context262K tokens (~131 books)
- ReleasedMar 2026
- LicenseProprietary
- StatusActive
- Cost / Message~$0.003
Frequently Asked Questions
Related Models
Qwen3 Coder Next FP8 · AlibabaQwen2.5 Coder 32B Instruct AWQ · AlibabaSpeaker Diarization Community 1 · PyannoteGemma 3 12B (free) · Google DeepMindQwen3 Next 80B A3B Instruct (free) · Alibaba QwenKey facts · as of 2026-05-03
- MiMo-V2-Omni by xiaomi. Not enough public benchmark scores to rank yet.
- List price $0.40 input · $2.00 output per 1M tokens (as of 2026-05-03).
How to cite · data as of 2026-05-03
MiMo-V2-Omni · benchmarks, pricing and providers. BenchGecko, data as of 2026-05-03. https://benchgecko.ai/model/mimo-v2-omni
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP