DeepSeek-R1-Distill-Llama-8B

8B

DeepSeek

R1 reasoning distilled into Llama 3.1 8B. Best chain-of-thought for 8–12GB cards; huge community GGUF support.

9.2K HF downloads53 likesbartowski/DeepSeek-R1-Distill-Llama-8B-GGUF· stats from 7/27/2026
Consumer GPUMac / Apple SiliconCPU / VPS

131K

Max Context

4

Quant Variants

GGUF Q5_K_M

Best Quality

98.8%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.855.6 GB2.6%145 tok/sCommunity
CalcHF
GGUFQ5_K_M5.686.5 GB1.2%128 tok/sEstimated
CalcHF
EXL24.65bpw4.655.3 GB1.9%210 tok/sEstimated
CalcHF
AWQINT444.9 GB3.5%188 tok/sEstimated
CalcHF