Qwen3 1.7B Instruct
1.7BAlibaba Qwen3
Tiny Qwen3 with thinking mode. Q4 ~1.4GB — phones, NUC, and always-on local agents.
Consumer GPUMac / Apple SiliconCPU / VPS
33K
Max Context
3
Quant Variants
GGUF Q8_0
Best Quality
99.6%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with Qwen3 8B8B
Qwen3 8B Instruct
Alibaba Qwen3
Consumer GPUMac / Apple Silicon
5.1 GBmin VRAM·99.4%accuracy
Latest Qwen3 dense 8B with thinking mode. Strong upgrade from Qwen2.5 7B for local deploy.
4B
Qwen3 4B Instruct
Alibaba Qwen3
Consumer GPUMac / Apple Silicon
2.8 GBmin VRAM·99.3%accuracy
Smallest Qwen3 dense with thinking mode. Q4 ~3.2GB — ideal for 8GB GPUs and edge devices.
14B
Qwen3 14B Instruct
Alibaba Qwen3
Consumer GPUMac / Apple Silicon
9.5 GBmin VRAM·98.8%accuracy
Qwen3 14B — best balance of reasoning and VRAM in the 2026 Qwen lineup.
30B-A3B
Qwen3 30B-A3B Instruct
Alibaba Qwen3
Consumer GPUMac / Apple Silicon
17.5 GBmin VRAM·98.9%accuracy
Qwen3 MoE with only 3B active params. Q4 ~19GB file; outperforms QwQ-32B on 16GB cards.