Context
131K tokens (~66 books)
Input $/1M
$0.34
Output $/1M
$0.34
Type
multimodal
License
Open Source
Benchmarks
0 tested
Data as of
About
Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and...
No benchmark data available yet.
Recently Happened
Llama 3.2 11B Vision Instruct pricing increased 400%
Apr 11, 2026
Links
Research
Documentation
Community
Source Code
BenchGecko API
llama-3-2-11b-vision-instruct
Specifications
- Typemultimodal
- Context131K tokens (~66 books)
- ReleasedSep 2024
- LicenseOpen Source
- StatusActive
- Cost / Message~$0.001
Available On
Share & Export
Frequently Asked Questions
Llama 3.2 11B Vision Instruct is an open-source multimodal AI model by Meta, released in September 2024. Context window: 131K tokens.
Related Models
Qwen3 Coder Next FP8 · AlibabaQwen2.5 Coder 32B Instruct AWQ · AlibabaSpeaker Diarization Community 1 · PyannoteGemma 3 12B (free) · Google DeepMindQwen3 Next 80B A3B Instruct (free) · Alibaba QwenBenchmarks
Key facts · as of 2026-07-11
- Llama 3.2 11B Vision Instruct by Meta. Not enough public benchmark scores to rank yet.
- List price $0.34 input · $0.34 output per 1M tokens (as of 2026-07-11).
How to cite · data as of 2026-07-11
Llama 3.2 11B Vision Instruct · benchmarks, pricing and providers. BenchGecko, data as of 2026-07-11. https://benchgecko.ai/model/llama-3-2-11b-vision-instruct
Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP