Qwen3 1.7B Instruct

1.7B

Alibaba Qwen3

Tiny Qwen3 with thinking mode. Q4 ~1.4GB — phones, NUC, and always-on local agents.

22.2K HF downloads55 likesQwen/Qwen3-1.7B-GGUF· stats from 7/27/2026
Consumer GPUMac / Apple SiliconCPU / VPS

33K

Max Context

3

Quant Variants

GGUF Q8_0

Best Quality

99.6%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.851.4 GB3.0%280 tok/sCommunity
CalcHF
GGUFQ8_08.52.1 GB0.4%240 tok/sEstimated
CalcHF
AWQINT441.2 GB4.0%320 tok/sEstimated
CalcHF

Similar models

Compare with Qwen3 8B