Llama 3.2 11B Vision Instruct
11BMeta Llama 3.2
Multimodal Llama with image understanding. Vision encoder adds ~2GB VRAM overhead.
Consumer GPUMac / Apple Silicon
131K
Max Context
2
Quant Variants
GGUF Q8_0
Best Quality
99.5%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Llama 3.23B
Llama 3.2 3B Instruct
Meta Llama 3.2
Consumer GPUMac / Apple Silicon
2.0 GBmin VRAM·99.8%accuracy
Tiny but capable. Runs on 4GB VRAM or 8GB RAM, even on phones via llama.cpp.
1B
Llama 3.2 1B Instruct
Meta Llama 3.2
Consumer GPUMac / Apple Silicon
1.0 GBmin VRAM·99.5%accuracy
Ultra-light Llama for mobile and embedded. Sub-2GB VRAM with Q4.
90B
Llama 3.2 90B Vision Instruct
Meta Llama 3.2
Pro GPU
44.2 GBmin VRAM·97.2%accuracy
Flagship multimodal Llama. Requires dual 4090 or A100; vision adds ~3GB overhead.
12B
Gemma 3 12B IT
Google Gemma 3
Consumer GPUMac / Apple Silicon
8.0 GBmin VRAM·98.7%accuracy
Mid-size Gemma 3 with vision. Fits 16GB at Q4; excellent multilingual chat.