Llama 4 Scout 17B (16E)
109B MoEMeta Llama 4
Meta Llama 4 Scout MoE (17B active / 109B total). Multimodal; needs ~68GB VRAM at Q4_K_M.
Pro GPU
10486K
Max Context
3
Quant Variants
GGUF Q4_K_M
Best Quality
97.6%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Llama 4400B MoE
Llama 4 Maverick 17B (128E)
Meta Llama 4
Pro GPU
198.0 GBmin VRAM·97.8%accuracy
Llama 4 Maverick flagship MoE (17B active / 400B total). Multi-GPU or H100 cluster territory.
675B MoE
Mistral Large 3 675B Instruct
Mistral AI
Pro GPU
312.0 GBmin VRAM·97.9%accuracy
Mistral 3 flagship MoE (41B active / 675B total) with vision encoder. FP8 on 8×H200; GGUF quant for research clusters only.
70B
Llama 3.1 70B Instruct
Meta Llama 3.1
Pro GPUMac / Apple Silicon
33.4 GBmin VRAM·98.8%accuracy
Meta's frontier 70B model. Requires 40GB+ VRAM; dual 3090 or M2 Ultra.
72B
Qwen2.5 72B Instruct
Alibaba Qwen2.5
Pro GPU
33.8 GBmin VRAM·98.9%accuracy
Flagship Qwen2.5. Requires dual 4090 or A100 80G. Exceptional reasoning at scale.