Qwen3-VL 30B-A3B Instruct

30B-A3B

Alibaba Qwen3-VL

Multimodal MoE with only ~3B active parameters, so it stays responsive on Apple unified memory and survives CPU offload far better than a dense 30B. Q4 ~19GB — a 24GB card holds it outright.

406.2K HF downloads591 likesQwen/Qwen3-VL-30B-A3B-Instruct· stats from 8/14/2026
Consumer GPUPro GPUMac / Apple Silicon

41K

Max Context

4

Quant Variants

GGUF Q8_0

Best Quality

99.7%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.8519.0 GB2.6%95 tok/sCommunity
CalcHF
GGUFQ3_K_M3.8715.2 GB5.4%104 tok/sEstimated
CalcHF
GGUFQ8_08.532.6 GB0.3%78 tok/sEstimated
CalcHF
AWQINT4417.2 GB3.7%112 tok/sEstimated
CalcHF