Home/Models/Llama 3.2 11B Vision Instruct
Meta logo

Llama 3.2 11B Vision Instruct

by Meta · Released Sep 2024

Open SourceMultimodal
Compare
Context
131K tokens (~66 books)
Input $/1M
$0.34
Output $/1M
$0.34
Type
multimodal
License
Open Source
Benchmarks
0 tested
Data as of
About

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and...

No benchmark data available yet.

Recently Happened
Llama 3.2 11B Vision Instruct pricing increased 400%
Apr 11, 2026
Links
Documentation
Community
BenchGecko API
llama-3-2-11b-vision-instruct
Specifications
  • Typemultimodal
  • Context131K tokens (~66 books)
  • ReleasedSep 2024
  • LicenseOpen Source
  • StatusActive
  • Cost / Message~$0.001
Available On
Meta logoMeta$0.34
Share & Export
Tweet
Llama 3.2 11B Vision Instruct is an open-source multimodal AI model by Meta, released in September 2024. Context window: 131K tokens.

Key facts · as of 2026-07-11

  • Llama 3.2 11B Vision Instruct by Meta. Not enough public benchmark scores to rank yet.
  • List price $0.34 input · $0.34 output per 1M tokens (as of 2026-07-11).

How to cite · data as of 2026-07-11

Llama 3.2 11B Vision Instruct · benchmarks, pricing and providers. BenchGecko, data as of 2026-07-11. https://benchgecko.ai/model/llama-3-2-11b-vision-instruct

Credit "Source: BenchGecko" with a link. Prices per provider and Gecko Tests are BenchGecko data (CC BY 4.0); benchmark scores keep their original source, listed in the JSON. JSON · llms.txt · MCP