8GB VRAM

RTX 3070 Ti — what LLMs can it run?

34 of 79 indexed models fit comfortably in 8GB at 4K context, each at the highest-quality quant that still leaves headroom.

14B · 5

RTX 3070 Ti — what LLMs can it run?14B
ModelQuantEst. VRAMHeadroomtok/s
Stable LM 2 12B Chat12BAWQ INT47.01 GB+1 GB142
Jamba 1.5 Mini12BAWQ INT46.82 GB+1.2 GB125
Solar 10.7B Instruct11BAWQ INT46.42 GB+1.6 GB168
Falcon 3 10B Instruct10BGPTQ INT46.12 GB+1.9 GB155
Gemma 2 9B Instruct9BAWQ INT46.27 GB+1.7 GB188

7B · 19

RTX 3070 Ti — what LLMs can it run?7B
ModelQuantEst. VRAMHeadroomtok/s
GLM-4-9B-Chat9BQ4_K_M5.87 GB+2.1 GB135
Qwen3-VL 8B Instruct8BQ4_K_M6.19 GB+1.8 GB140
Qwen2-VL 7B Instruct7BQ4_K_M5.49 GB+2.5 GB72
Granite 3.1 8B Instruct8BQ4_K_M5.88 GB+2.1 GB142
Qwen3 8B Instruct8BEXL2 4.65bpw5.59 GB+2.4 GB228
Llama 3.1 8B Instruct8BEXL2 4.65bpw5.43 GB+2.6 GB235
Nous Hermes 3 Llama 3.1 8B8BEXL2 4.65bpw5.43 GB+2.6 GB232
Aya 23 8B8BQ4_K_M5.64 GB+2.4 GB145
OpenChat 3.6 8B8BEXL2 4.65bpw5.43 GB+2.6 GB228
DeepSeek-R1-Distill-Llama-8B8BQ5_K_M6.51 GB+1.5 GB128
InternLM2 7B Chat7BQ4_K_M5.45 GB+2.6 GB148
Qwen2.5 7B Instruct7BQ6_K6.77 GB+1.2 GB132
Qwen2.5-Coder 7B Instruct7BEXL2 4.65bpw4.87 GB+3.1 GB248
WizardLM-2 7B7BQ4_K_M5.07 GB+2.9 GB152
DeepSeek-R1-Distill-Qwen-7B7BEXL2 4.65bpw4.87 GB+3.1 GB210
OLMo 2 7B Instruct7BQ4_K_M6.82 GB+1.2 GB150
Mistral 7B Instruct v0.37BQ6_K6.75 GB+1.3 GB135
Zephyr 7B Beta7BQ6_K6.75 GB+1.3 GB132
Gemma 3 4B IT4BQ8_05.36 GB+2.6 GB145

≤3B · 10

RTX 3070 Ti — what LLMs can it run?≤3B
ModelQuantEst. VRAMHeadroomtok/s
Qwen3 4B Instruct4BQ6_K4.05 GB+4 GB145
Phi-4 Mini Instruct3.8BQ8_04.68 GB+3.3 GB262
Phi-3.5 Mini Instruct3.8BQ8_05.89 GB+2.1 GB255
Llama 3.2 3B Instruct3BQ8_04.05 GB+4 GB285
Qwen2.5 3B Instruct3BQ8_03.67 GB+4.3 GB290
Gemma 2 2B Instruct2BQ8_03.34 GB+4.7 GB320
Qwen3 1.7B Instruct1.7BQ8_02.39 GB+5.6 GB240
Qwen2.5 1.5B Instruct1.5BQ8_01.83 GB+6.2 GB410
Llama 3.2 1B Instruct1BQ8_01.51 GB+6.5 GB450
Qwen2.5 0.5B Instruct0.5BQ8_00.6 GB+7.4 GB540

How this list is built

Each row is the lowest-perplexity-loss quant of that model whose estimated total — weights plus KV cache at 4K context plus activation buffer — uses at most 88% of the card. That is the calculator's "green" threshold, so every row here has real headroom rather than only just fitting. Raise the context length and the list shortens; the calculator lets you check any combination directly.

Stepping up

A RTX 3080 10G (10GB) fits 9 more of the indexed models than this card. RTX 3080 10G

34 of 79 indexed models fit comfortably in 8GB at 4K context, each at the highest-quality quant that still leaves headroom.