DeepSeek-R1-Distill-Llama-8B
8BDeepSeek
R1 reasoning distilled into Llama 3.1 8B. Best chain-of-thought for 8–12GB cards; huge community GGUF support.
Consumer GPUMac / Apple SiliconCPU / VPS
131K
Max Context
4
Quant Variants
GGUF Q5_K_M
Best Quality
98.8%
Accuracy Retained
Quantization Variants
Per-quant VRAM, quality loss, and inference speed on RTX 4090
Measured = site benchmarks · Estimated = formula · Community = public reports
Similar models
Compare with DeepSeek-R1-Distill-Qwen-7B7B
DeepSeek-R1-Distill-Qwen-7B
DeepSeek
Consumer GPUMac / Apple Silicon
5.2 GBmin VRAM·97.8%accuracy
R1 reasoning in a 7B footprint. Best value for 8–12GB VRAM CoT experiments.
16B
DeepSeek-V2-Lite Chat
DeepSeek
Consumer GPUMac / Apple Silicon
9.6 GBmin VRAM·97.0%accuracy
MoE general model (~2.4B active). Long context and strong multilingual chat.
14B
DeepSeek-R1-Distill-Qwen-14B
DeepSeek
Consumer GPU
9.2 GBmin VRAM·98.0%accuracy
R1 reasoning distilled into 14B. Huge community interest; excellent chain-of-thought.
32B
DeepSeek-R1-Distill-Qwen-32B
DeepSeek
Consumer GPUPro GPU
16.8 GBmin VRAM·97.4%accuracy
R1 distilled to 32B. Near-frontier reasoning on a single 24GB card (Q3/Q4).