12GB VRAM

RTX 5070 — what LLMs can it run?

46 of 81 indexed models fit comfortably in 12GB at 4K context, each at the highest-quality quant that still leaves headroom.

14B · 16

RTX 5070 — what LLMs can it run?14B
ModelQuantEst. VRAMHeadroomtok/s
DeepSeek-Coder-V2-Lite Instruct16BAWQ INT49.13 GB+2.9 GB192
DeepSeek-V2-Lite Chat16BAWQ INT49.13 GB+2.9 GB188
StarCoder2 15B15BQ4_K_M10.19 GB+1.8 GB92
Qwen3 14B Instruct14BEXL2 4.65bpw9.66 GB+2.3 GB125
Qwen2.5 14B Instruct14BEXL2 4.65bpw9.75 GB+2.3 GB138
DeepSeek-R1-Distill-Qwen-14B14BEXL2 4.65bpw9.75 GB+2.3 GB128
Phi-4 14B14BQ4_K_M10.17 GB+1.8 GB88
Phi-3 Medium 14B Instruct14BQ4_K_M9.73 GB+2.3 GB102
Mistral Nemo 12B Instruct12BQ4_K_M8.42 GB+3.6 GB112
Gemma 3 12B IT12BQ4_K_M9.28 GB+2.7 GB105
Stable LM 2 12B Chat12BQ4_K_M8.35 GB+3.7 GB108
Jamba 1.5 Mini12BQ4_K_M8.15 GB+3.9 GB95
Llama 3.2 11B Vision Instruct11BQ4_K_M7.66 GB+4.3 GB88
Solar 10.7B Instruct11BQ4_K_M7.6 GB+4.4 GB125
Falcon 3 10B Instruct10BQ4_K_M7.28 GB+4.7 GB118
Gemma 2 9B Instruct9BQ4_K_M7.3 GB+4.7 GB132

7B · 20

RTX 5070 — what LLMs can it run?7B
ModelQuantEst. VRAMHeadroomtok/s
GLM-4-9B-Chat9BQ8_010.16 GB+1.8 GB105
Qwen3-VL 8B Instruct8BQ8_010.39 GB+1.6 GB108
Ministral 3 8B Instruct8BQ8_010.02 GB+2 GB
Qwen2-VL 7B Instruct7BQ4_K_M5.49 GB+6.5 GB72
Granite 3.1 8B Instruct8BQ4_K_M5.88 GB+6.1 GB142
Qwen3 8B Instruct8BQ6_K7.64 GB+4.4 GB122
Llama 3.1 8B Instruct8BQ8_09.47 GB+2.5 GB118
Nous Hermes 3 Llama 3.1 8B8BEXL2 4.65bpw5.43 GB+6.6 GB232
Aya 23 8B8BQ4_K_M5.64 GB+6.4 GB145
OpenChat 3.6 8B8BEXL2 4.65bpw5.43 GB+6.6 GB228
DeepSeek-R1-Distill-Llama-8B8BQ5_K_M6.51 GB+5.5 GB128
InternLM2 7B Chat7BQ4_K_M5.45 GB+6.6 GB148
Qwen2.5 7B Instruct7BQ6_K6.77 GB+5.2 GB132
Qwen2.5-Coder 7B Instruct7BEXL2 4.65bpw4.87 GB+7.1 GB248
WizardLM-2 7B7BQ4_K_M5.07 GB+6.9 GB152
DeepSeek-R1-Distill-Qwen-7B7BEXL2 4.65bpw4.87 GB+7.1 GB210
OLMo 2 7B Instruct7BQ8_010.3 GB+1.7 GB125
Mistral 7B Instruct v0.37BQ6_K6.75 GB+5.3 GB135
Zephyr 7B Beta7BQ6_K6.75 GB+5.3 GB132
Gemma 3 4B IT4BQ8_05.36 GB+6.6 GB145

≤3B · 10

RTX 5070 — what LLMs can it run?≤3B
ModelQuantEst. VRAMHeadroomtok/s
Qwen3 4B Instruct4BQ6_K4.05 GB+8 GB145
Phi-4 Mini Instruct3.8BQ8_04.68 GB+7.3 GB262
Phi-3.5 Mini Instruct3.8BQ8_05.89 GB+6.1 GB255
Llama 3.2 3B Instruct3BQ8_04.05 GB+8 GB285
Qwen2.5 3B Instruct3BQ8_03.67 GB+8.3 GB290
Gemma 2 2B Instruct2BQ8_03.34 GB+8.7 GB320
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+9.6 GB240
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+10.2 GB410
Llama 3.2 1B Instruct1BQ8_01.51 GB+10.5 GB450
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+11.4 GB540

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

A further 2 models load but with no headroom to spare (up to 105% of VRAM) — the Quant Hub’s GPU chips count those too, which is why its number is higher.

Measured on this card

No benchmark runs in this index were recorded on a RTX 5070. Every figure on this page is calculated from the model architecture and the quant level — treat them as estimates, not measurements.

Cards with the same budget

What fits is decided by memory, so every 12GB card of this type returns the same list. These pages are not different answers — they differ in throughput, which this index does not measure per card.

Stepping up

A RTX 5080 (16GB) fits 6 more of the indexed models than this card. RTX 5080

46 of 81 indexed models fit comfortably in 12GB at 4K context, each at the highest-quality quant that still leaves headroom.