VRAM / RAM Calculator
Precise memory estimate for any model × quant × context combination
Pick a model and see if it fits your hardware
KV cache type
f16 is the default in llama.cpp and Ollama.
Select a model and format above to see memory estimate
Quick reference by model size
The calculator's answer for every size class in this index, at each model's reference quant (its native build, else Q4_K_M) and batch 1 — the same arithmetic as the tool above, written out so it can be read without it. The 32K column counts only models whose own context window reaches 32K. The last column is the smallest card that holds every model in the class comfortably at 4K.
| Size class | Models | At 4K context | At 32K context | Holds the whole class |
|---|---|---|---|---|
| ≤3B | 10 | 0.4–4.1 GB | 0.7–15.6 GB | Mac M1 8G |
| 7B | 15 | 3.0–6.8 GB | 3.2–10.9 GB | RTX 4060 (8 GB) |
| 14B | 11 | 7.3–10.2 GB | 10.4–15.9 GB | RTX 3060 12G |
| 32B | 21 | 12.8–30.4 GB | 13.5–31.6 GB | A100 40G |
| 70B+ | 14 | 46.1–428.7 GB | 55.7–436.4 GB | No single card — split or offload |
How the estimate is calculated
Three terms are added. Model weights are params × bpw ÷ 8, plus 2% for embedding and norm tensors that most quantizers keep at higher precision. KV cache is 2 (K and V) × layers × kv_heads × head_dim × context × batch × bytes per value: 2 for the default fp16 cache, and ggml's block sizes when you pick a quantized one — q8_0 stores 32 values in 34 bytes, q4_0 in 18. Activation buffer is a flat 10% of the first two, covering the transient tensors an inference runtime allocates per forward pass. Units: every figure is computed and displayed in binary gigabytes (GiB, 1024³ bytes) but labelled "GB", matching how GPU vendors and nvidia-smi label VRAM. There is no conversion between the two anywhere in the tool. When you pick an indexed model the calculator uses that model's own measured bits-per-weight rather than the generic per-level table, because the two can differ sharply — GPT-OSS ships mostly-MXFP4 weights, so its Q8_0 is 5.10 bpw, not 8.5.
Why context length matters more than you expect
Weights are fixed once you pick a quant level; the KV cache is not. It grows linearly with context and with batch size, and on long-context models it can overtake the weights entirely. The lever that decides how steep that growth is is grouped-query attention: kv_heads is often far smaller than the number of attention heads. Llama 3.1 8B has 32 attention heads but only 8 KV heads, so its cache is a quarter of what multi-head attention would cost. Two models of identical size can therefore have very different memory curves, which is why this calculator reads layers, kvHeads and headDim from each model's real config rather than assuming a shape.
How close is it to reality
For Llama 3.1 8B at Q4_K_M the estimate is 4.62 GB against a 4.58 GiB file from bartowski — close enough to plan a card purchase around. Treat it as an estimate all the same. Runtimes differ in how much they reserve: llama.cpp's compute buffer scales with batch and context, vLLM pre-allocates a fixed fraction of the card (--gpu-memory-utilization, 0.85 by default in our generated commands), and CUDA itself takes a few hundred MB of context before any weights load. The verdict colours build that in — green means the estimate uses at most 88% of the card, amber up to 105%, so a green result already has headroom.
Common questions
- Does this include the operating system and display overhead?
- No. It estimates what the model needs. On a card that is also driving a desktop, subtract roughly 0.5–1.5 GB before comparing — which is part of why the green threshold stops at 88% rather than 100%.
- Why does the same quant level give different sizes for different models?
- Because bits-per-weight is an average over a mixed scheme.
Q4_K_Mkeeps some tensors at higher precision, and how many depends on the architecture — an MoE model with most of its parameters in experts quantizes differently from a dense one. Where a model in the index ships that level, the calculator uses its measured figure instead of the generic one. - Can I run a model that shows amber?
- Often yes, with less context, a smaller batch or a quantized KV cache — all three shrink the cache directly. With your card chosen, the calculator prints the longest context that stays green and the longest that stays amber. If it is still amber at short context, the weights alone are the problem and you need a smaller quant level; the table of every level the model ships shows which one fits.
- What if the model is bigger than my card?
- Two ways, and the calculator sizes both. Add a card: llama.cpp's default split gives each card a slice of layers and that slice's KV cache, so memory adds up — Llama 3.3 70B at Q4_K_M and 8K context is 47.5 GB, out of reach for one RTX 3090 and a tight fit on two. The cards take turns on every token, so a second card buys room, not speed. Spill to system RAM: llama.cpp and Ollama keep as many whole layers on the GPU as fit and run the rest on the CPU. On one RTX 3090 that is about 38 of the 80 layers, with about 25 GB in system RAM, and the RAM half sets the pace — the bound is roughly 4 tok/s with dual-channel DDR5-5600. For a mixture-of-experts model,
--n-cpu-moemoves only the expert weights, which usually costs far less speed. - Should I quantize the KV cache?
- At long context it is the cheapest memory you can win back. Llama 3.1 8B at Q4_K_M and 32K context is 9.49 GB with the default fp16 cache, 7.42 GB with q8_0 and 6.32 GB with q4_0 — the weights do not change, only the cache does. In llama.cpp pass
-ctk q8_0 -ctv q8_0; a quantized V cache needs Flash Attention, which-fa autoturns on and which fails if forced off. In Ollama setOLLAMA_KV_CACHE_TYPE=q8_0with Flash Attention enabled. Ollama's own documentation describes q8_0 as usually unnoticeable and q4_0 as a small-to-medium precision loss that shows more at long context; this site has not measured either. - What does reverse mode do differently?
- Forward mode answers "will this model fit on my card". Reverse mode starts from the card and lists every model × quant pair in the index that fits, sorted by quality, speed or footprint. Both size from the same per-model figures, so the two views agree.
- Are these numbers measured or estimated?
- The VRAM figure is computed, not measured. The per-quant
vramGBand speed values in the model index carry a confidence marker — measured, estimated or community — and the benchmarks page documents the hardware and framework versions behind the measured ones.