Qwen3-VL 8B Instruct
8BAlibaba Qwen3-VL
Current-generation vision-language model that still fits a single 8–12GB card at Q4 (~5.9GB). The realistic multimodal option for people without a 24GB GPU — note the vision encoder adds VRAM the KV-cache math below does not model.
41K
Max Context
4
Quant Variants
GGUF Q8_0
Best Quality
99.7%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3-VL 30B-A3BQwen3-VL 30B-A3B Instruct
Alibaba Qwen3-VL
Multimodal MoE with only ~3B active parameters, so it stays responsive on Apple unified memory and survives CPU offload far better than a dense 30B. Q4 ~19GB — a 24GB card holds it outright.
Gemma 3 12B IT
Google Gemma 3
Mid-size Gemma 3 with vision. Fits 16GB at Q4; excellent multilingual chat.
Qwen2.5 14B Instruct
Alibaba Qwen2.5
The sweet spot between performance and resource usage. 16GB VRAM with Q4.
Mistral Nemo 12B Instruct
Mistral AI
Mistral + NVIDIA collaboration. 128K context, excellent multilingual support.
How to actually run this
Deployment guides for this model and this class of hardware.