Llama 3.2 90B Vision Instruct
90BMeta Llama 3.2
Flagship multimodal Llama. Requires dual 4090 or A100; vision adds ~3GB overhead.
Pro GPU
131K
Max Context
2
Quant Variants
GGUF Q4_K_M
Best Quality
97.2%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Llama 3.211B
Llama 3.2 11B Vision Instruct
Meta Llama 3.2
Consumer GPUMac / Apple Silicon
9.5 GBmin VRAM·99.5%accuracy
Multimodal Llama with image understanding. Vision encoder adds ~2GB VRAM overhead.
3B
Llama 3.2 3B Instruct
Meta Llama 3.2
Consumer GPUMac / Apple Silicon
2.0 GBmin VRAM·99.8%accuracy
Tiny but capable. Runs on 4GB VRAM or 8GB RAM, even on phones via llama.cpp.
1B
Llama 3.2 1B Instruct
Meta Llama 3.2
Consumer GPUMac / Apple Silicon
1.0 GBmin VRAM·99.5%accuracy
Ultra-light Llama for mobile and embedded. Sub-2GB VRAM with Q4.
109B MoE
Llama 4 Scout 17B (16E)
Meta Llama 4
Pro GPU
55.0 GBmin VRAM·97.6%accuracy
Meta Llama 4 Scout MoE (17B active / 109B total). Multimodal; needs ~68GB VRAM at Q4_K_M.