DeepSeek-R1-Distill-Llama-8B

8B

DeepSeek

R1 reasoning distilled into Llama 3.1 8B. Best chain-of-thought for 8–12GB cards; huge community GGUF support.

12.3K HF downloads53 likesbartowski/DeepSeek-R1-Distill-Llama-8B-GGUF· stats from 9/13/2026
Consumer GPUMac / Apple SiliconCPU / VPS

131K

Max Context

4

Quant Variants

GGUF Q5_K_M

Best Quality

98.8%

Accuracy Retained

Quantization Variants

Per-quant VRAM, quality loss, and inference speed on RTX 4090

Measured = site benchmarks · Estimated = formula · Community = public reports

FormatLevelBPWVRAMPPL LossSpeedSourceActions
GGUFQ4_K_M4.855.6 GB2.6%145 tok/sCommunity
CalcHF
GGUFQ5_K_M5.686.5 GB1.2%128 tok/sEstimated
CalcHF
EXL24.65bpw4.655.3 GB1.9%210 tok/sEstimated
CalcHF
AWQINT444.9 GB3.5%188 tok/sEstimated
CalcHF

Running DeepSeek-R1-Distill-Llama-8B locally

At Q4_K_M and 4K of context, DeepSeek-R1-Distill-Llama-8B needs about 5.6 GB — 4.6 GB of weights, 0.50 GB of KV cache and a 0.5 GB activation buffer. The smallest card in this index that clears that comfortably is the RTX 5060 Ti 8G at 8 GB, and 61 of the 61 cards here do. These are calculated figures, not measurements: the estimate stops counting a card as comfortable at 88% of its VRAM, which is roughly the room a desktop session needs.

What longer context costs

Going from 4K to 32K adds about 3.50 GB, taking the total to 9.5 GB. Weights do not move with context — only the KV cache does, and it grows linearly, so this is the number to watch when planning for long documents. The model's native window is 128K; holding all of it at Q4_K_M would need about 23 GB.

Which build to download

This index tracks 3 formats for it — GGUF, EXL2, AWQ — across 4 levels. Q5_K_M carries the lowest published perplexity loss at 1.2%. The fastest level measured here is EXL2 4.65bpw at 210 tok/s on an RTX 4090, batch 1. GGUF runs on llama.cpp and Ollama across NVIDIA, AMD and Apple silicon; AWQ and GPTQ target vLLM on CUDA and ROCm; EXL2 is ExLlamaV2 and CUDA only.

Common questions

How much VRAM does DeepSeek-R1-Distill-Llama-8B need?
About 5.6 GB at Q4_K_M with 4K of context and batch 1, rising to roughly 9.5 GB at 32K. That figure is weights plus KV cache plus a 10% activation buffer, calculated from the model's architecture rather than measured on a card.
Will DeepSeek-R1-Distill-Llama-8B run on a 8GB GPU?
Yes — at Q4_K_M and 4K context it needs about 5.6 GB, which leaves 2.4 GB spare on a RTX 5060 Ti 8G. That is the smallest card in this index that clears it comfortably; 61 of 61 do.
Which quantization of DeepSeek-R1-Distill-Llama-8B should I use?
Q5_K_M has the lowest published quality loss (1.2%), and Q4_K_M is the level most people run. All 4 levels in the index are Q4_K_M, Q5_K_M, EXL2 4.65bpw, AWQ INT4.

Where this model fits

Sized at Q4_K_M with a 4K context window, smallest card first. Comfortable means the estimate uses at most 88% of the memory.

Ships in
GGUFEXL2AWQ

This page's figures change when the model or the runtime does.Last updated 2026-09-13 RSS → /feed.xml