Qwen2.5-Coder 7B Instruct

7B

Alibaba Qwen2.5

Best 7B coding model. Ideal for local dev assistants on 8–16GB VRAM.

176.3K HF downloads366 likesQwen/Qwen2.5-Coder-7B-Instruct-GGUF· stats from 8/8/2026
Consumer GPUMac / Apple SiliconCPU / VPS

131K

Max Context

3

Quant Variants

EXL2 4.65bpw

Best Quality

98.0%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.855.4 GB2.8%158 tok/sEstimated
CalcHF
AWQINT444.8 GB4.0%225 tok/sEstimated
CalcHF
EXL24.65bpw4.655.2 GB2.0%248 tok/sEstimated
CalcHF